Papers for 2026-09-29

10 papers
Tianyu Xiong, Yi Lu, Jinrui Wang, ... , Qiu Shen, Xun Cao
总结: 提出BeyondRetarget端到端框架,直接从单目RGB视频学习可执行的人形机器人动作,绕过显式人体表示与运动重定向。
方法: 抛弃以人体表示为中心的级联流程,直接学习面向机器人的隐式表示,把视频映射为机器人动作并联合优化,避免人体估计误差传播。
证据: 相比先估计人体动作再重定向的传统管线,该方法显著提升动作在机器人上的可执行性与精度,论文报告了仿真与实机层面的对比验证。
为什么适合我: 为我提供了绕过运动重定向的新范式,可直接从海量人视频获取可跟踪的参考动作,契合运动模仿数据规模化与迁移到真实机器人的需求。
原摘要

Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.

Zizhuo Wang, Ming-ju Lee, Shaoting Zhu, ... , Hang Zhao, Yiming Li
总结: 提出TactileStep,将足底压力触觉引入人形行走控制,实现更柔软的落地与更稳定的支撑接触,应对跑酷式苛刻足地交互。
方法: 对齐触觉仿真与真实压力鞋垫特征,利用触觉与运动线索识别足接触相位,并施加相位感知奖励引导安全落地与稳定站立。
证据: 在仿真和Unitree G1人形实机上评估,表明触觉反馈能改善硬着陆、边缘接触与不稳定支撑等任务指标无法反映的接触质量问题。
为什么适合我: 触觉是接触丰富控制的关键模态,该工作为我提供了把足底触觉融入全身策略、缩小人机感知差距并可实机部署的具体路径。
原摘要

Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a key domain gap between humans and humanoid robots: the absence of rich tactile sensing in most humanoid systems. We address this problem with TactileStep, a deployable tactile learning framework that brings sole pressure sensing into humanoid locomotion control for softer touchdowns and more stable support. TactileStep aligns tactile simulation with the real pressure insole, allowing the policy to learn from the same contact features available on hardware. During training, we use tactile and motion cues to recognize different foot-contact phases and apply phase-aware rewards that encourage safer landing and more stable stance. Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.

Xukun Luan, Zhongxiang Lei, Chen Gong, ... , Yuanguo Bi, Jinyan Liu
总结: 提出首个面向物理人形控制的动作级遗忘方法ForgetMimic,可从RL策略中选择性删除特定动作而保留其余运动能力。
方法: 给定在N个动作上训练的策略,通过遗忘算法降低其在K个目标动作上的性能,同时保持其余N-K个动作的执行效果,并解决两个关键失效问题。
证据: 以安全、隐私与GDPR被遗忘权为动机,在物理人形控制任务上验证了目标动作性能退化与其余动作保持的双重效果。
为什么适合我: 运动模仿策略的动作数据涉及版权与隐私,遗忘技术为我管理和净化参考动作库、合规部署生成式运动策略提供了新工具。
原摘要

Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains insufficiently explored. Addressing this issue is motivated by pressing safety and privacy concerns: the removal of malicious, poisoned, or suboptimal motions, as well as copyright-protected motions subject to the right to be forgotten under regulations such as the GDPR, is of critical importance. To this end, we propose {ForgetMimic}, the first motion-level unlearning method designed specifically for physical-world humanoid control. The core idea of ForgetMimic is as follows: given a policy $π_θ$ trained on $N$ motions, our method degrades performance on a target subset of $K$ motions while preserving the effectiveness of the remaining $N-K$ motions. Furthermore, we identify and resolve two key training mechanisms in robot control that lead to unlearning failure. We conduct extensive experiments on the Unitree G1 and H2 humanoid robots across 12 motions, including Dance, Fight, Flip, and others. Experimental results demonstrate that ForgetMimic effectively eliminates memory of designated motions while maintaining the normal operation of all other motions.

Kyrylo Kolesnichenko, Irvin Steve Cardenas, Jong-Hoon Kim
总结: 提出端到端框架,将单目T台走秀视频转化为可部署的人形行走策略,使人形机器人复现有表现力的猫步与时尚风格。
方法: 串联动作恢复、机器人重定向、动作校正、策略训练、仿真评估与实机部署的完整管线,把视频风格化步态转化为全身控制策略。
证据: 在Booster K1人形机器人上完成全部实机试验且无一跌倒,复现了特征性窄步距及腿、躯干、手臂的协调运动。
为什么适合我: 展示了从单目视频到全身运动跟踪再到实机部署的完整闭环,对我构建视频到实机的动作模仿与风格化步态管线有直接参考价值。
原摘要

Runway walking requires coordinated control of posture, stride, foot placement, and whole-body motion to effectively present clothing and convey a distinctive style. However, humanoid robots used in fashion shows typically rely on locomotion policies optimized primarily for stability and walking speed, limiting their ability to reproduce expressive, human-like runway motions. In this work, we present an end-to-end framework that transforms monocular runway videos into deployable humanoid locomotion policies through motion recovery, robot retargeting, motion correction, policy training, simulation-based evaluation, and physical deployment. We evaluate the proposed framework on the Booster K1 humanoid robot using runway-style catwalk motions. The learned policy completed every physical trial without falling, while reproducing the characteristic narrow foot placement and coordinated movement of the legs, torso, and arms. The results demonstrate that our proposed training framework enables the Booster K1 to perform stable and expressive catwalk motions, highlighting its potential for humanoid robotic applications in fashion shows and other performance-oriented scenarios.

Zepeng Wang, Jiangxing Wang, Chao Ma, Xiaochuan Shi, Zongqing Lu
总结: 提出PLAT,让人形策略仅凭稀疏定时关键帧即可执行稳定全身运动,衔接密集动作跟踪与稀疏目标条件控制两种范式。
方法: 三阶段框架:预训练密集跟踪专家提供运动先验,经DAgger式模仿学习特权潜变量先验,再用潜变量残差强化学习精调状态转移。
证据: 训练时以密集目标序列作特权监督,部署时只需稀疏定时关键帧,在稳定到达连续运动目标上优于直接稀疏控制基线。
为什么适合我: 稀疏关键帧接口让跟踪策略可充当规划与交互式运动生成的高级控制器,正契合我构建通用可跟踪全身控制器的目标。
原摘要

Humanoid motion tracking policies rely on dense frame-by-frame references, limiting their use as high-level motion controllers for planning and interactive motion generation. We study \emph{Sparse Timed Keyframe Motion Tracking}, where a policy receives only sparse future keyframes and their desired arrival times, and must execute stable whole-body motions that reach successive goals. We propose \textbf{PLAT}, a three-stage sparse timed keyframe motion tracking policy learning framework with \textbf{P}rivileged \textbf{LA}tent \textbf{T}ransition learning. PLAT bridges dense motion tracking and sparse goal-conditioned control by exploiting dense goal sequences as privileged supervision during training while requiring only sparse timed keyframe commands at deployment. A pretrained dense tracking expert first provides robust motion priors. A privileged latent prior is then learned through DAgger-style imitation, followed by latent residual reinforcement learning that refines latent transitions instead of directly optimizing actions. Extensive simulation experiments demonstrate that PLAT maintains accurate and stable sparse timed keyframe tracking across varying planning horizons, with particularly strong performance under long-horizon commands. Successful deployment on a Unitree G1 humanoid robot further demonstrates the effectiveness and practicality of PLAT for sparse humanoid motion control.

Seoyeon Choi, Shizhao Ye, Nicholas Bui, ... , Markus Wulfmeier, Negar Mehr
总结: 提出HuGo,用大语言模型在冻结的低层全身策略之上生成可执行的高级策略代码,实现人形移动操作任务的免训练泛化。
方法: LLM根据任务、观测与命令规范从任务描述生成闭环高级策略代码,再利用数值轨迹与精选视频帧反馈做针对性代码更新与迭代。
证据: 在五个仿真任务、两种不同低层全身策略上验证,HuGo显著优于依赖奖励工程或演示加任务训练的基线方法。
为什么适合我: 提供免逐任务训练的分层新思路,LLM代码化任务逻辑可复用我已有的全身低层控制器,快速扩展移动操作技能库。
原摘要

For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/

Toan Nguyen, Weiduo Yuan, Siheng Zhao, Yue Wang, Daniel Seita
总结: 提出HOTICE全身学习框架,使人形机器人在机器人与负载周围空间均受限的杂乱环境中协同避障地搬运物体。
方法: 设计人形-物体解耦势场联合编码双方避障引导,双智能体RL解耦上下肢控制并经共享观测与奖励保持协调,再用专家到通才蒸馏泛化。
证据: 在多样化杂乱场景中训练,特权教师蒸馏出的通才策略可泛化到未见场景,实现全身协调的物体搬运。
为什么适合我: 上下肢解耦加共享奖励的架构直接回应我关心的全身协调难题,势场引导与蒸馏策略对接触丰富的移动操作很有设计参考价值。
原摘要

Object transportation is a fundamental capability for humanoid robots operating in real-world, human-centric environments, yet existing methods struggle when clutter constrains free space around both the robot and its carried payload. We present HOTICE, a whole-body humanoid learning framework for transporting objects through such cluttered environments. First, we introduce Humanoid-Object Decoupled Potential Fields, which jointly encode collision-avoidance guidance for the robot and the carried object, enabling coordinated, obstacle-aware motion for both. Second, to address the large action space inherent to whole-body loco-manipulation, we design a dual-agent reinforcement learning architecture that decouples upper- and lower-body control while preserving whole-body coordination via shared state observations and rewards. To train a policy that generalizes across diverse cluttered scenes, we further employ a specialist-to-generalist distillation strategy, in which privileged teacher policies are distilled into a single deployable student policy. We evaluate HOTICE in MuJoCo simulation and on a real Unitree G1 humanoid, demonstrating effective and robust object transportation across cluttered scenarios for objects of varying shapes. Our results show that HOTICE reliably coordinates whole-body motion and object-aware collision avoidance, generalizing effectively to previously unseen cluttered environments while achieving strong performance in sim2real deployment.

Zachary Olkin, William D. Compton, Aaron D. Ames
总结: 提出两层感知运动架构:流匹配生成器从深度图像规划全身轨迹,感知跟踪策略执行,并通过离策略RL微调持续改进生成器。
方法: 用动力学优化的人类数据构建地形一致动作库训练双策略,以结构化搜索采集数据,经优势加权回归离策略微调流匹配运动生成器。
证据: 该离策略微调环比在线残差微调样本效率显著更高,在未见地形几何与技能组合上提升地形一致性与速度跟踪精度。
为什么适合我: 生成式运动先验加感知跟踪正是我构建多技能人形控制器的核心组件,其样本高效的生成器微调环路尤其值得借鉴。
原摘要

General purpose humanoids require locomotion controllers that are multi-skill, perceptive, dynamic, and robust enough to go anywhere humans can. In this work, we present a two layer locomotion architecture: (1) a perceptive flow matching motion generator plans whole body trajectories from raw depth images while a (2) perceptive tracking policy trained with control-guided RL follows these motions. Both policies are trained on a library of terrain consistent motion clips created with dynamically optimized human data which yields both accurate velocity tracking and terrain consistent references. Our central contribution is a simple yet effective off-policy RL fine tuning loop that improves the motion generator. A structured search method is used with the generator to gather data for advantage weighted regression. This off-policy loop is much more sample efficient than on-policy residual fine tuning and improves terrain consistency on unseen geometries and skill compositions. We find that successful terrain traversals increased by up to 25 percentage points and skill selection improved by up to 80 percentage points. By using raw depth images to perceive the environment no odometry or height maps are needed, and outdoor deployment is easy. With two cameras, the policy can see terrain coming from further away and adjust its velocity regardless of the commanded speed so it can traverse the terrain. A single policy pair enables a Unitree G1 humanoid to walk, run, stand, jump on and off of boxes, and traverse stairs in outdoor environments. Project page: https://zolkin1.github.io/generate-track-improve/

Yichuan Yu, Youzhuo Wang, Yiming Ren, ... , Yujing Sun, Yuexin Ma
总结: 提出MATE多智能体虚拟遥操作平台,让多地操作员在共享物理仿真中同时控制多台人形机器人,规模化采集协作数据。
方法: 多个分布式操作员在共享物理环境中控制全身人形,保留人形、物体与环境间的物理耦合交互,免除多台实机、专用场地与反复复位。
证据: 构建了24.1小时、2500条联合片段、五个长程任务的多智能体人形协作数据集,涵盖物体交接、接力运送与环境交互等。
为什么适合我: 为我提供了低成本规模化采集多人形协作与移动操作数据的途径,可补充真机遥操作数据以扩展模仿学习训练集。
原摘要

Humanoid robots require diverse embodied experiences to acquire complex loco-manipulation and collaborative skills. However, existing humanoid data pipelines primarily focus on individual agents, while physical multi-robot collaboration remains difficult to scale due to costly hardware, dedicated spaces, and repeated resets. In this work, we introduce MATE, a Multi-Agent virtual TEleoperation platform for humanoid collaboration data collection that enables multiple geographically distributed operators to simultaneously control whole-body humanoids in a shared physics-based environment. MATE removes the need for multiple physical robots and co-located operation while preserving physically coupled interactions among humanoids, objects, and environments. Using MATE, we construct a multi-humanoid collaboration dataset comprising 24.1 hours of coordinated behavior across 2,500 joint episodes and five long-horizon tasks, including object handover, relay delivery, environment interaction, and cooperative transport. To improve learning from these interaction-rich demonstrations, we introduce EAIS, an Execution-Aligned Interaction Sampling strategy that computes sampling signals within an execution-aligned prefix and prioritizes task-progressing and interaction-critical behaviors. We evaluate MATE with representative imitation learning and vision-language-action policies across diverse collaboration tasks. Experiments demonstrate efficient data collection, effective policy learning, and zero-shot transfer from virtual demonstrations to a physical humanoid without real-world fine-tuning. Project page: https://yerik-yu.github.io/MATE/

Yuanzhu Zhan, Yufei Jiang, Zemu Zhang, Junyi Geng
总结: 提出户外真实环境飞行操作框架,融合模仿学习策略、机载LiDAR惯性状态估计与全身MPC,实现野外无动捕的飞行操作。
方法: 用扩散策略从演示学习并从机载观测生成期望末端运动,全身MPC联合协调飞行平台与机械臂跟踪该轨迹,机载LiDAR惯性里程计提供状态估计。
证据: 在真实飞行操作平台上完成户外操作任务验证,全程无需外部动捕基础设施,展示了抗扰动的全身协调执行。
为什么适合我: 虽聚焦飞行平台,但其学习策略输出轨迹、全身MPC跟踪的分层架构,对我设计人形感知运动与全身控制的结合很有启发。
原摘要

Aerial manipulation in outdoor environments remains challenging due to the simultaneous requirements of reliable state estimation, stable aerial motion, and precise manipulation under external disturbances. In this work, we present a real-world outdoor aerial manipulation framework that integrates imitation learning, onboard LiDAR-inertial state estimation, and whole-body model predictive control. A Diffusion Policy is trained from manipulation demonstrations to generate desired end-effector motions from onboard observations. These learned commands are executed by a whole-body MPC that jointly coordinates the aerial platform and manipulator to realize the desired end-effector trajectory. To eliminate reliance on external motion-capture infrastructure, the platform employs onboard LiDAR-inertial odometry for state estimation during outdoor operation. We validate the complete framework on a physical aerial manipulator and demonstrate successful execution of outdoor manipulation tasks. The experimental results show that demonstration-driven manipulation policies can be effectively integrated with onboard state estimation and model-based whole-body control to enable aerial manipulation beyond controlled indoor environments.