Papers for 2026-03-06

10 papers
Beichen Wang, Yuanjie Lu, Linji Wang, Liuchuan Yu, Xuesu Xiao
总结: 本文提出开源VR数据采集与评测框架MTC,用于在可控杂乱3D场景中采集并重定向全身人类运动到类人机器人,构建包含145个场景348条轨迹的数据集,并提供衡量杂乱程度、稳定性和碰撞安全性的基准,以推动场景感知类人机器人运动研究。
原摘要

Recent advances in humanoid locomotion have enabled dynamic behaviors such as dancing, martial arts, and parkour, yet these capabilities are predominantly demonstrated in open, flat, and obstacle-free settings. In contrast, real-world environments such as homes, offices, and public spaces, are densely cluttered, three-dimensional, and geometrically constrained, requiring scene-aware whole-body coordination, precise balance control, and reasoning over spatial constraints imposed by furniture and household objects. However, humanoid locomotion in cluttered 3D environments remains underexplored, and no public dataset systematically couples full-body human locomotion with the scene geometry that shapes it. To address this gap, we present Moving Through Clutter (MTC), an opensource Virtual Reality (VR) based data collection and evaluation framework for scene-aware humanoid locomotion in cluttered environments. Our system procedurally generates scenes with controllable clutter levels and captures embodiment-consistent, whole-body human motion through immersive VR navigation, which is then automatically retargeted to a humanoid robot model. We further introduce benchmarks that quantify environment clutter level and locomotion performance, including stability and collision safety. Using this framework, we compile a dataset of 348 trajectories across 145 diverse 3D cluttered scenes. The dataset provides a foundation for studying geometry-induced adaptation in humanoid locomotion and developing scene-aware planning and control methods.

Sijia Li, Haoyu Wang, Shenghai Yuan, Yizhuo Yang, Thien-Minh Nguyen
总结: 本文提出一种四足机器人导航的分层策略TDGC,将高层基于稀疏地形线索的任务决策与低层强化学习步态控制解耦并通过可调接口连接,从而提升混合地形和分布外环境中的稳定性、可迁移性与任务成功率。
原摘要

Real-world quadruped navigation is constrained by a scale mismatch between high-level navigation decisions and low-level gait execution, as well as by instabilities under out-of-distribution environmental changes. Such variations challenge sim-to-real transfer and can trigger falls when policies lack explicit interfaces for adaptation. In this paper, we present a hierarchical policy architecture for quadrupedal navigation, termed Task-level Decision to Gait Control (TDGC). A low-level policy, trained with reinforcement learning in simulation, delivers gait-conditioned locomotion and maps task requirements to a compact set of controllable behavior parameters, enabling robust mode generation and smooth switching. A high-level policy makes task-centric decisions from sparse semantic or geometric terrain cues and translates them into low-level targets, forming a traceable decision pipeline without dense maps or high-resolution terrain reconstruction. Different from end-to-end approaches, our architecture provides explicit interfaces for deployment-time tuning, fault diagnosis, and policy refinement. We introduce a structured curriculum with performance-driven progression that expands environmental difficulty and disturbance ranges. Experiments show higher task success rates on mixed terrains and out-of-distribution tests.

Weikai Qin, Sichen Wu, Ci Chen, ... , Haoqi Han, Hesheng Wang
总结: PhysiFlow提出一种融合语义运动意图、物理约束与多脑潜在流匹配的类人机器人全身VLA控制框架,实现了更高效且稳定的视觉语言引导全身协调执行。
原摘要

In the domain of humanoid robot control, the fusion of Vision-Language-Action (VLA) with whole-body control is essential for semantically guided execution of real-world tasks. However, existing methods encounter challenges in terms of low VLA inference efficiency or an absence of effective semantic guidance for whole-body control, resulting in instability in dynamic limb-coordinated tasks. To bridge this gap, we present a semantic-motion intent guided, physics-aware multi-brain VLA framework for humanoid whole-body control. A series of experiments was conducted to evaluate the performance of the proposed framework. The experimental results demonstrated that the framework enabled reliable vision-language-guided full-body coordination for humanoid robots.

Mingzhe Li, Mengyin Liu, Zekai Wu, ... , Siqi Shen, Cheng Wang
总结: 本文提出“运动图灵测试”来仅基于运动学评估人形机器人动作的人类相似度,构建含人类与11种机器人动作的HHMotion数据集并发现机器人在跳跃、拳击、跑步等动态动作上仍明显不像人,同时给出自动评分基准且简单模型优于现有多模态大模型方法。
原摘要

Humanoid robots have achieved significant progress in motion generation and control, exhibiting movements that appear increasingly natural and human-like. Inspired by the Turing Test, we propose the Motion Turing Test, a framework that evaluates whether human observers can discriminate between humanoid robot and human poses using only kinematic information. To facilitate this evaluation, we present the Human-Humanoid Motion (HHMotion) dataset, which consists of 1,000 motion sequences spanning 15 action categories, performed by 11 humanoid models and 10 human subjects. All motion sequences are converted into SMPL-X representations to eliminate the influence of visual appearance. We recruited 30 annotators to rate the human-likeness of each pose on a 0-5 scale, resulting in over 500 hours of annotation. Analysis of the collected data reveals that humanoid motions still exhibit noticeable deviations from human movements, particularly in dynamic actions such as jumping, boxing, and running. Building on HHMotion, we formulate a human-likeness evaluation task that aims to automatically predict human-likeness scores from motion data. Despite recent progress in multimodal large language models, we find that they remain inadequate for assessing motion human-likeness. To address this, we propose a simple baseline model and demonstrate that it outperforms several contemporary LLM-based methods. The dataset, code, and benchmark will be publicly released to support future research in the community.

Valerii Serpiva, Jeffrin Sam, Chidera Simon, ... , Artem Lykov, Dzmitry Tsetserukou
总结: DreamToNav通过将自然语言指令转化为视觉描述并利用生成式视频模型“预演”机器人行为,从合成视频中提取可执行轨迹,实现了无需任务特定工程、可跨轮式与四足平台泛化的室内导航。
原摘要

We present DreamToNav, a novel autonomous robot framework that uses generative video models to enable intuitive, human-in-the-loop control. Instead of relying on rigid waypoint navigation, users provide natural language prompts (e.g. ``Follow the person carefully''), which the system translates into executable motion. Our pipeline first employs Qwen 2.5-VL-7B-Instruct to refine vague user instructions into precise visual descriptions. These descriptions condition NVIDIA Cosmos 2.5, a state-of-the-art video foundation model, to synthesize a physically consistent video sequence of the robot performing the task. From this synthetic video, we extract a valid kinematic path using visual pose estimation, robot detection and trajectory recovery. By treating video generation as a planning engine, DreamToNav allows robots to visually "dream" complex behaviors before executing them, providing a unified framework for obstacle avoidance and goal-directed navigation without task-specific engineering. We evaluate the approach on both a wheeled mobile robot and a quadruped robot in indoor navigation tasks. DreamToNav achieves a success rate of 76.7%, with final goal errors typically within 0.05-0.10 m and trajectory tracking errors below 0.15 m. These results demonstrate that trajectories extracted from generative video predictions can be reliably executed on physical robots across different locomotion platforms.

Duncan Andrews, Landon Zimmerman, Evan Martin, Joe DiGennaro, Baxi Chong
总结: 该论文提出一种小型仿蜥蜴机器人SILA Bot,利用本体感觉估计颗粒地形深度并通过简单线性反馈调节身体波动模式,从而在未知复杂基质上实现高效自适应运动。
原摘要

Unlike their large-scale counterparts, small-scale robots are largely confined to laboratory environments and are rarely deployed in real-world settings. As robot size decreases, robot-terrain interactions fundamentally change; however, there remains a lack of systematic understanding of what sensory information small-scale robots should acquire and how they should respond when traversing complex natural terrains. To address these challenges, we develop a Small-scale, Intelligent, Lizard-inspired, Adaptive Robot (SILA Bot) capable of adapting to diverse substrates. We use granular media of varying depths as a controlled yet representative terrain paradigm. We show that the optimal body movement pattern (ranging from standing-wave bending that assists limb retraction on flat ground to traveling-wave undulation that generates thrust in deep granular media) can be parameterized and approximated as a linear function of granular depth. Furthermore, proprioceptive signals, such as joint torque, provide sufficient information to estimate granular depth via a K-Nearest Neighbors classifier, achieving 95% accuracy. Leveraging these relationships, we design a simple linear feedback controller that modulates body phase and substantially improves locomotion performance on terrains with unknown depth. Together, these results establish a principled framework for perception and control in small-scale locomotion and enable effective terrain-adaptive locomotion while maintaining low computational complexity.

Arnau Boix-Granell, Alberto San-Miguel-Tello, Magí Dalmau-Moreno, Néstor García
总结: PRISM 将模仿学习与强化学习结合,并利用自然语言指令和人类反馈迭代细化奖励与策略,使机器人能在少量数据下将通用操作技能高效适配到新的目标和约束,在模拟抓取放置任务中提升了鲁棒性并降低计算成本。
原摘要

This paper presents PRISM: an instruction-conditioned refinement method for imitation policies in robotic manipulation. This approach bridges Imitation Learning (IL) and Reinforcement Learning (RL) frameworks into a seamless pipeline, such that an imitation policy on a broad generic task, generated from a set of user-guided demonstrations, can be refined through reinforcement to generate new unseen fine-grain behaviours. The refinement process follows the Eureka paradigm, where reward functions for RL are iteratively generated from an initial natural-language task description. Presented approach, builds on top of this mechanism to adapt a refined IL policy of a generic task to new goal configurations and the introduction of constraints by adding also human feedback correction on intermediate rollouts, enabling policy reusability and therefore data efficiency. Results for a pick-and-place task in a simulated scenario show that proposed method outperforms policies without human feedback, improving robustness on deployment and reducing computational burden.

Ziken Huang, Xinze Niu, Bowen Chai, Renbiao Jin, Danping Zou
总结: Swooper 通过两阶段深度强化学习训练单个轻量级策略网络,在简单夹爪四旋翼上实现零样本高速空中抓取,真实实验中以最高 1.5 m/s 速度达到 84% 成功率。
原摘要

High-speed aerial grasping presents significant challenges due to the high demands on precise, responsive flight control and coordinated gripper manipulation. In this work, we propose Swooper, a deep reinforcement learning (DRL) based approach that achieves both precise flight control and active gripper control using a single lightweight neural network policy. Training such a policy directly via DRL is nontrivial due to the complexity of coordinating flight and grasping. To address this, we adopt a two-stage learning strategy: we first pre-train a flight control policy, and then fine-tune it to acquire grasping skills. With the carefully designed reward functions and training framework, the entire training process completes in under 60 minutes on a standard desktop with an Nvidia RTX 3060 GPU. To validate the trained policy in the real world, we develop a lightweight quadrotor grasping platform equipped with a simple off-the-shelf gripper, and deploy the policy in a zero-shot manner on the onboard Raspberry Pi 4B computer, where each inference takes only about 1.0 ms. In 25 real-world trials, our policy achieves an 84% grasp success rate and grasping speeds of up to 1.5 m/s without any fine-tuning. This matches the robustness and agility of state-of-the-art classical systems with sophisticated grippers, highlighting the capability of DRL for learning a robust control policy that seamlessly integrates high-speed flight and grasping. The supplementary video is available for more results. Video: https://zikenhuang.github.io/Swooper/.

TADPO: Reinforcement Learning Goes Off-road

4.0 Candidate 当日相对
Zhouchonghao Wu, Raymond Song, Vedant Mundheda, ... , Christof Schoenborn, Jeff Schneider
总结: TADPO通过结合离策略教师轨迹指导与在线学生探索来扩展PPO,实现了视觉端到端强化学习在复杂越野地形中的高速驾驶,并首次在全尺寸越野车上展示了零样本仿真到现实迁移。
原摘要

Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics. Addressing these challenges requires effective long-horizon planning and adaptable control. Reinforcement Learning (RL) offers a promising solution by learning control policies directly from interaction. However, because off-road driving is a long-horizon task with low-signal rewards, standard RL methods are challenging to apply in this setting. We introduce TADPO, a novel policy gradient formulation that extends Proximal Policy Optimization (PPO), leveraging off-policy trajectories for teacher guidance and on-policy trajectories for student exploration. Building on this, we develop a vision-based, end-to-end RL system for high-speed off-road driving, capable of navigating extreme slopes and obstacle-rich terrain. We demonstrate our performance in simulation and, importantly, zero-shot sim-to-real transfer on a full-scale off-road vehicle. To our knowledge, this work represents the first deployment of RL-based policies on a full-scale off-road platform.

Pei Qu, Zheng Li, Yufei Jia, ... , Jinni Zhou, Jun Ma
总结: OmniDP利用全景LiDAR点云和时间感知注意力池化实现端到端3D视觉运动策略,使类人机器人无需频繁移动底座即可在大范围、杂乱环境中进行稳健操作,并在仿真和真实实验中优于传统自视角深度相机方法。
原摘要

The deployment of humanoid robots for dexterous manipulation in unstructured environments remains challenging due to perceptual limitations that constrain the effective workspace. In scenarios where physical constraints prevent the robot from repositioning itself, maintaining omnidirectional awareness becomes far more critical than color or semantic information.While recent advances in visuomotor policy learning have improved manipulation capabilities, conventional RGB-D solutions suffer from narrow fields of view (FOV) and self-occlusion, requiring frequent base movements that introduce motion uncertainty and safety risks. Existing approaches to expanding perception, including active vision systems and third-view cameras, introduce mechanical complexity, calibration dependencies, and latency that hinder reliable real-time performance. In this work, We propose OmniDP, an end-to-end LiDAR-driven 3D visuomotor policy that enables robust manipulation in large workspaces. Our method processes panoramic point clouds through a Time-Aware Attention Pooling mechanism, efficiently encoding sparse 3D data while capturing temporal dependencies. This 360° perception allows the robot to interact with objects across wide areas without frequent repositioning. To support policy learning, we develop a whole-body teleoperation system for efficient data collection on full-body coordination. Extensive experiments in simulation and real-world environments show that OmniDP achieves robust performance in large-workspace and cluttered scenarios, outperforming baselines that rely on egocentric depth cameras.