Papers for 2026-10-03

10 papers
Jaeryeong Kim, Taerim Yoon, Jin Cheng, Sungjoon Choi, Stelian Coros
总结: 提出密集时序运动重定向方法DTMR,将人类动作适配到足式机器人动态特性,尤其改善跳跃等时序敏感的动态动作。
方法: 在单次优化中联合调整时序与控制,每个控制步都进行时间变形,用GPU并行采样模型预测控制求解。
证据: 在两小时人类动作数据、四种人形机器人上评估,DTMR全面优于基线方法,动态动作上优势最明显。
为什么适合我: 与我的运动重定向和全身技能学习直接相关,密集时序变形思路可提升动态参考动作向真实机器人的迁移质量。
原摘要

Legged robots can learn expressive whole-body skills from the motions of humans and animals. Due to the morphology gap between the source and the robot, however, the motion must be tailored to the dynamic properties of the robot. In particular, dynamic motions such as a jump require careful adjustment, since their timing and control are interdependent. We propose dense temporal motion retargeting (DTMR), which jointly optimizes timing and control within a single optimization, where dense means that the timing is adjusted for every control step. This dense formulation enables DTMR to deform only the parts of the motion that need a change in timing. The problem is solved with sampling-based model predictive control (MPC) in parallel on a GPU. We evaluate DTMR against baselines on two hours of human motion with four humanoid robots, where the results show that DTMR outperforms baseline methods, particularly on dynamic motions. We also show that allowing more temporal deformation yields more precise retargeting. We further compare DTMR with a baseline that optimizes the temporal dimension, where the result shows that DTMR retargets more precisely under the same deformation budget while being ~19x faster. Lastly, policies trained on our references transfer to a real humanoid robot.

Chongyang Xu, Zhao Wu, Jin Chen, ... , Li Lu, Steven C. H. Hoi
总结: 用500小时自我中心人类全身数据预训练,构建通用人形移动操作视觉动作模型,探索以人类经验支撑可扩展学习。
方法: 轻便可穿戴系统同步采集自我中心视频与身体手部动作,构建HumanVerse-500数据集并训练全身视觉策略λ0。
证据: 数据集覆盖开放世界中多样移动操作行为,无需机器人操作即可提供大规模全身运动与手物协调监督。
为什么适合我: 为我的感知运动与全身控制器研究提供可扩展的人类数据来源,有望缓解遥操作数据采集难以扩展的瓶颈。
原摘要

Humanoid whole-body manipulation has advanced rapidly, enabling policies to coordinate locomotion, posture, bimanual interaction, and dexterous hand movements. Meanwhile, egocentric human videos provide diverse examples of everyday interactions across objects and scenes, offering scalable supervision without robot operation. However, existing supervision from these videos provides limited coverage of whole-body movement and coordination with hand-object interaction, while obtaining such supervision through humanoid teleoperation is also costly and difficult to scale. We therefore explore how human experience can support scalable learning of humanoid loco-manipulation. To support this study, we introduce HumanVerse-500, a 500-hour dataset of diverse human loco-manipulation behaviors in open-world environments, collected with a lightweight wearable system that synchronizes egocentric video with body and hand motion. Building on this dataset, we develop $λ_0$, a whole-body humanoid vision-language-action policy, through three-stage training that first learns interaction from diverse egocentric datasets, then coordinates body and hand motion using HumanVerse-500, and finally adapts the policy to downstream tasks and robot embodiments. Across these stages, $λ_0$ learns a shared representation space for human experience transfer, while domain-specific interfaces handle differences between human and robot states and actions. We evaluate $λ_0$ on SIMPLE and 4 real-world loco-manipulation tasks, achieving state-of-the-art performance, and further analyze its scaling behavior, generalization, and training-stage contributions to understand how human data support downstream whole-body humanoid control. We will release our code, models, and data to support further research.

Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger
总结: 将四足步态中命令跟踪、稳定性与能效的权衡从固定奖励变为运行时可调的偏好输入,实现单一可部署策略。
方法: 偏好条件化多目标强化学习,策略以部署侧偏好为输入,固定与本体相关的运动先验,分离操作意图与奖励塑形。
证据: 仿真中采样100个偏好,67种行为在精确Pareto支配下非支配,优于固定目标、多目标基线与独立专门策略。
为什么适合我: 单策略可调目标权衡便于真实部署,可借鉴到我的四足双足控制器中统一管理多种任务优先级而无需重训。
原摘要

Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.

Kuankuan Sima, Yichao Gao, Chenxi Gu, Kefan Zhao, Lin Zhao
总结: 通过响应一致性的运动策略与策略感知MPC耦合,实现足式机器人边行走边精确跟踪末端目标。
方法: 响应塑形使策略在随机动力学下命令响应一致可重复,再辨识闭环响应模型,供MPC联合规划运动命令与手臂动作。
证据: 仿真基准上位置与姿态RMSE较各自最佳基线分别降低28.7%和27.4%,并展示了真实世界实验效果。
为什么适合我: RL底层加MPC上层正是我关注的移动操作架构,响应一致性建模使学习策略可预测,便于与上层规划器协同。
原摘要

Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.

Yifan Hu, Luhang Hong, Mingkang Long, ... , Junjie Fu, Guanghui Wen
总结: 提出去中心化多智能体强化学习框架,在复用预训练单机器人技能之上学习高层策略,实现多人形全身协调移动操作。
方法: 共享去中心化高层策略调用预训练技能,仅需任务级奖励;对同构多人形引入置换数据增强并证明其保持策略梯度方向。
证据: 理论上证明置换样本不改变策略梯度方向,实验表明无需任务特定动作参考或大量奖励工程即可学到协调行为。
为什么适合我: 技能复用与去中心化协调可把我的全身运动控制扩展到多人形协作场景,置换增强思路也适用于同构多机器人训练。
原摘要

Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.

Abu Hanif Muhammad Syarubany, Jaehyun Jang, Hwanhee Kim, ... , Seungyeon Ryu, Chang D. Yoo
总结: 颠倒人形足球系统堆栈:先训练通用命令条件运动基座,再以任务门控层叠加多方向踢球技能库,运动本身成为独立能力。
方法: 命令条件运动策略作基座,N个运动引导踢球技能作为门控层,技能从可命令状态出发并返回,转换数从O(N²)降至O(N)。
证据: 在29自由度Unitree G1上实现七种重定向踢球技能,覆盖259.5度名义方向,踢后稳定交还运动控制器处理。
为什么适合我: 运动基座加技能门控的组合范式契合我的全身控制器设计理念,可让敏捷步态与多样技能解耦并按需组合部署。
原摘要

Recent humanoid soccer systems make motion tracking the substrate and derive locomotion from it, typically by steering a motion-reference anchor toward the ball. This yields strong shooting results, but locomotion is trained only on the narrow, deterministic command distribution ball approach induces, never evaluated as a capability in its own right. We invert the stack: a general, command-conditioned locomotion policy is trained first as the substrate, and N motion-guided kicking skills are added on top as task-gated layers, so the reachable gait space is set by the locomotion curriculum rather than any reference clip. Because every skill starts from and returns to this same commandable state, locomotion also becomes a composition hub (O(N) transitions rather than O(N^2)), and post-strike stabilisation is handed back to the trained controller rather than scripted per clip. We instantiate this on a 29-DoF Unitree G1 with seven retargeted kicking skills spanning 259.5 degrees of nominal aim direction, including lateral, rearward and weak-foot strikes a single forward-facing reference cannot express, and report shooting accuracy alongside command-tracking, terrain and push-recovery results with the full skill library attached, an axis prior humanoid soccer systems do not report. The library is validated on hardware across forward, lateral, rearward and commanded approaches.

Xinyuan Luo, Chunyuan Yang, Boyuan Chen, Xianyi Cheng
总结: 提出人形移动操作柔顺框架,实现末端定向可调柔顺与躯干可选柔顺,在接触外力时兼顾任务运动精度。
方法: 分层RL控制器通过高层末端与躯干命令调制固定全身跟踪策略,交互力由本体感知历史估计得出。
证据: 仿真与真实实验验证同一末端可在一个方向顺应接触力、另一方向保持精度,躯干可拒力或随力运动。
为什么适合我: 接触丰富环境需要力响应能力,把柔顺叠加在跟踪策略之上的分层思路契合我的真实机器人全身控制迁移目标。
原摘要

Humanoids are increasingly capable of tracking complex whole-body motions, but physical interaction introduces a different challenge. When a robot makes contact with a person or the environment, it needs to respond to external forces while preserving the motion needed for the task. This response can vary across directions in the end-effectors and on the body. For example, an end effector may need to accommodate contact force in one direction while maintaining motion accuracy in another, while the robot body may resist an external force or move with it. We present a compliance framework for humanoid loco-manipulation that combines directional and tunable end-effector (EE) compliance with selectable root compliance for external force rejection or force following. A hierarchical reinforcement learning controller modulates a fixed whole-body tracking policy through high-level EE and root commands, while interaction forces are estimated from proprioceptive history. Our simulation and real-world experiments on a humanoid demonstrate directional stiffness control, online stiffness adjustment, distinct root compliance, compliant manipulation, and collaborative carrying.

Xiangyu Miao, Junsong Wu, Jiyuan Shi, ... , Chenjia Bai, Xuelong Li
总结: 提出感知全身控制框架NEXUS,使遥操作人形在与操作者地形不一致时自适应调整姿态与接触,而非逐帧复制动作。
方法: 可扩展地形感知适配算法自动生成跨动作与地形的高质量配对数据,近1000小时语料经师生学习训练感知全身控制器。
证据: 生成近1000小时配对动作语料,无需逐动作逐地形调参即可训练出跨地形自适应的遥操作全身控制器。
为什么适合我: 遥操作与地形自适应正处我的研究交叉点,其配对数据生成方法可缓解感知运动全身控制学习的监督稀缺问题。
原摘要

Whole-body teleoperation requires a humanoid robot to reproduce a human operator's behavior even when their terrains differ. This demands that the robot perceive local terrain and adapt its posture and contacts accordingly, rather than copy the operator's motion frame by frame. However, paired motion data linking the same behaviors across flat ground and different terrains remain scarce, limiting supervision for learning terrain-adaptive control. To enable whole-body teleoperation across mismatched terrains, we introduce NEXUS, a perceptive whole-body control framework that combines human motion commands with onboard sensory feedback. We first develop a scalable terrain-aware adaptation algorithm that efficiently generates high-quality motion pairs across motions and terrains without per-motion or per-terrain tuning. Using a paired motion corpus totaling nearly 1,000 hours, we train a perceptive whole-body controller through teacher-student learning to reproduce commanded behaviors across terrains. Experiments demonstrate efficient, scalable generation of high-quality motion data and show that NEXUS combines broad behavioral coverage with terrain adaptability and tracking fidelity, outperforming existing whole-body controllers on the evaluated benchmarks. Zero-shot real-world deployment enables real-time whole-body teleoperation on diverse unseen terrains, further validating the generalization of our method. Project website: https://nexus-humanoid.github.io/

Yoshiki Takebayashi, Giovanni Perantoni, Hikaru Sasaki, Matteo Saveriano, Takamitsu Matsubara
总结: 提出EF-GAIfO,从机器人自身经验估计仅状态示范的可行性,解决本体差异导致人类示范对机器人不可行的问题。
方法: 基于生成对抗的从观测模仿,可行域由机器人经验估计并随策略改进逐步扩展,动态纳入更多示范,无需动力学模型。
证据: 在无动作标签示范设置下,可行性感知与渐进扩展机制提升了策略性能,已在仿真任务中得到验证。
为什么适合我: 契合我的运动重定向与生成式策略方向,可行性动态估计可用于筛选人类动作参考,改善仿真到真实的迁移效果。
原摘要

With the increasing use of robot-free demonstration interfaces that provide state trajectories without action labels, imitation from observation has become a promising approach for learning robot behaviors from human demonstrations. However, due to differences in embodiment and dynamics between humans and robots, demonstrated human motions may not be feasible for the robot, potentially degrading policy performance. In this study, we propose Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), which estimates the feasibility of state-only demonstrations from the robot's own experience rather than relying on explicit dynamics models or large prior exploration datasets. A key feature of EF-GAIfO is that the notion of feasibility evolves with policy learning: as the policy improves and the robot experiences a broader range of state transitions, the feasible region is progressively expanded, allowing additional demonstrations to be incorporated into learning. This enables feasibility-aware imitation that adapts to the current stage of policy learning, rather than relying on a pre-designed feasibility criterion. We validate the effectiveness of EF-GAIfO on a locomotion task in simulation and on a real quadruped robot performing a object-reaching-and-grasping task.

Yin Gu, Xinming Zhang, Shanze Wang, Siwei Cheng, Wei Zhang
总结: 提出程序化分层导航框架LME,通过显式生成与推理子目标,引导移动机器人在未知环境中逃离局部极小区域。
方法: 仅用局部观测,以可解释启发式准则结合障碍几何与候选位置安全性选择子目标,局部规划器生成底层运动命令。
证据: 在存在与不存在局部极小的未知环境中均稳健导航,恢复机制显式可解释而非依赖深度RL奖励隐式学习。
为什么适合我: 与我核心方向关联较弱,但其显式子目标与程序化恢复机制可借鉴到腿式机器人在非结构化地形导航的分层控制中。
原摘要

Mapless navigation in unknown and partially observable environments remains challenging for mobile robots, particularly when local minima prevent the robot from making progress toward its goal. Existing local navigation methods often lack an explicit mechanism for escaping such situations, while deep reinforcement learning (DRL) approaches typically learn recovery behaviors implicitly through reward design and policy optimization. In this work, we propose \textbf{LME} (Local-Minimum Escaper), a programmatic hierarchical framework that explicitly generates and reasons subgoals to guide robots out of local-minimum regions. LME operates solely on local observations and selects candidate subgoals using interpretable heuristic criteria that account for both surrounding obstacle geometry and candidate-location safety. A local planner then generates low-level motion commands toward the selected subgoal. This design enables LME to handle environments both with and without local minima within a unified framework, while remaining independent of the underlying local planner and requiring no additional training. Extensive experiments in simulated and real-world environments demonstrate that LME provides robust navigation performance and generalizes to challenging unseen scenarios. Furthermore, the generated subgoals can be used to guide different local planners, substantially improving their ability to escape local minima. Successful deployments on both differential-drive and quadruped robots further demonstrate the practical applicability and generality of the proposed framework.