Papers for 2026-09-23

10 papers
Utsav Panchal, Denis Kleyko, Unal Artan, Amy Loutfi
总结: 解耦上下身平滑约束的强化学习,实现人形稳定全身控制。
方法: 将平滑性作为上下身物理极限显式约束,并用有界障碍惩罚。
证据: 提升边界约束满足,兼顾任务响应与物理平滑。
为什么适合我: 契合人形全身控制中差异化平滑与稳定性需求。
原摘要

Embodied AI systems, particularly humanoid robots deployed in real world scenarios require whole-body control policies that are both task-responsive and physically smooth. However, smoothness is not uniform across the body: lower body must remain sufficiently reactive, while the upper body must be tightly regulated to preserve stability. Existing reinforcement learning approaches typically impose smoothness through auxiliary terms in the reward function, which compete with task objectives, treating the body as uniform and provide no direct control over the physical quantities responsible for smooth behavior. We introduce DeCap (Decoupled Constraint-aware policy), a constrained reinforcement learning algorithm that decouples whole-body smoothness into separate upper- and lower-body constraint groups, each formulates smoothness as explicit constraints on physical motion limits. To improve constraint satisfaction near feasibility boundaries, DeCap incorporates a bounded barrier penalty that activates proactively as limits are approached and remains bounded at the constraint limit. On real-world humanoid whole-body control task, DeCap reduces upper-body action rate by 2.50x and acceleration by 2.18x relative to reward-based smoothness policies, while also improving lower-body smoothness and reducing transient motion. We demonstrate that a fixed set of smoothness constraints transfers across diverse terrains, alleviating the need of extensive reward tuning.

Fukang Liu, Yipu Chen, Jaehwi Jang, ... , Zsolt Kira, Ye Zhao
总结: 力感知VLA框架,为人形接触丰富操作引入显式力命令。
方法: 多任务VLA联合预测几何运动目标与连续接触力参考。
证据: 解决接触后视觉不可靠及不同力制度需求。
为什么适合我: 直接支持接触丰富非结构化环境中的移动操作。
原摘要

Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions through geometric motion goals and rely on whole-body controllers focused on motion tracking, with limited explicit reasoning or control of interaction forces. This limitation is particularly relevant in contact-rich tasks, where geometrically similar motions may require different force regimes depending on the task context and where visual observations may become unreliable after contact. In this work, we present Opt2VLA, a force-aware VLA framework that introduces explicit force commands at the VLA-to-control interface for humanoid whole-body manipulation. A single multi-task VLA policy jointly predicts both geometric motion goals and continuous contact-force references, which are tracked by task-specific reinforcement learning (RL)-based whole-body controllers. To provide scalable and physically grounded supervision, we generate dynamically feasible and contact-consistent training data via whole-body trajectory optimization (TO) with explicit force references. We evaluate Opt2VLA on three contact-rich humanoid tasks and show that explicit force conditioning enables more accurate and consistent force regulation than motion-only control, while physically grounded torque supervision from TO further improves force tracking accuracy and stability. Closed-loop evaluations further demonstrate language-conditioned force modulation with Opt2VLA in simulation and on humanoid hardware.

Yuanzhuo Li, Wen Zhao, Zhe Yong, ... , Zhen Wang, Yijie Guo
总结: 分层多步态框架,实现人形精确落脚的3D loco-manipulation。
方法: 整合地形感知步进、AMP行走与上身控制,用LD-PPO蒸馏。
证据: 融合行走与步进专家为统一可执行学生策略。
为什么适合我: 利于敏捷行走、地形接触利用与全身操作结合。
原摘要

Humanoid loco-manipulation faces two prominent limitations: controllers using continuous velocity commands cannot precisely regulate individual footholds, while specialized foothold-tracking modules are difficult to integrate with whole-body manipulation. Furthermore, standard action-based imitation distillation primarily transfers expert actions, without explicitly encouraging a shared representation of heterogeneous skills. This paper introduces STRIDER, a hierarchical multi-gait framework to bridge these gaps. The framework integrates terrain-aware 3D stepping logic, Adversarial Motion Priors (AMP)-based natural walking, and Cartesian upper-body control: its stepping expert selects feasible footholds in the stance-foot frame and generates clearance-aware swing trajectories. To fuse distinct walking and stepping experts into one executable student policy, we propose Latent Distillation Proximal Policy Optimization (LD-PPO), a distillation algorithm augmented with teacher-conditioned latent alignment. By jointly optimizing on-policy reinforcement learning, DAgger-based action reconstruction, and latent alignment, LD-PPO transfers expert actions while encouraging a shared skill representation across heterogeneous modes. Simulation and real-robot evaluations on the TianGong Omni humanoid show that LD-PPO outperforms vanilla distillation-PPO in foothold-tracking and posture-tracking accuracy. Deployed on hardware, STRIDER realizes multi-gait loco-manipulation with accurate foothold and end-effector tracking.

Sicen Li, Zhen Chu, Chao Li, Qiuguo Zhu, Jun Wu
总结: 多源点级融合框架,支持人形挑战地形全身运动。
方法: LiDAR与深度相机早期融合为点集,线性自注意力编码。
证据: 保留薄障碍,单模态失效仅损失部分信息。
为什么适合我: 增强非结构化地形感知运动的覆盖与鲁棒性。
原摘要

Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360° light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.

Xu Han, Angsong Li, Shaopeng Zhang, ... , Yuan Zhuang, Haiyu Lan
总结: 先验知情里程计,从人体运动跟踪提升人形本体感知。
方法: 重定向人体运动生成监督,结合物理对称先验结构化预测。
证据: 真实机器人下相对最强基线误差降低31.6%-61.7%。
为什么适合我: 改善仿真到真实迁移,支持多样运动跟踪。
原摘要

Simulation-trained humanoid proprioceptive odometry faces two transfer challenges: training trajectories generated by specific robot control policies intended for deployment cover only a limited range of motions, while sim-to-real mismatch can make unconstrained predictions unreliable. We address both with Prior-Informed Odometry from Human-Motion Tracking (PRIMO). On the data side, we generate odometry supervision by having the humanoid track diverse retargeted human motions in simulation, decoupling supervision from the deployment policies and broadening the training motion distribution. On the model side, a Prior-Informed estimator uses physics- and symmetry-informed priors to structure velocity and rotation prediction and a coarse raw-context pathway to preserve sensor context alongside encoded features, thereby strengthening sim-to-real generalization. Under a unified real-robot protocol, PRIMO reduces mean error by 31.6%-61.7% relative to the strongest evaluated external baseline in each domain-metric comparison. Across two locomotion-policy revisions, policy specialists exhibit symmetric crossover, whereas Tracking-Locomotion training reduces mean opposite-policy simulation error by 86.8%-94.6%. On real dynamic motion, Tracking-Locomotion training reduces mean error by 69.2%-81.7% relative to training on the union of both deployment policies. Across the tested motion compositions, the Prior-Informed estimator consistently lowers mean trajectory errors relative to its Unconstrained counterpart in both simulation and real-robot evaluation. Code is available at https://github.com/Agibot-Spatial-Intelligence/PRIMO.

Yuxuan Nai, Leixin Chang, Liangjing Yang, Shuo Yang, Zhongyu Li
总结: 将UMI技能迁移至人形全身操作的实时运动生成。
方法: 端效应器条件生成器解耦协调,扩散策略学UMI演示。
证据: 异步层级整合策略生成器与控制器,含延迟补偿。
为什么适合我: 可扩展遥操作与全身移动操作数据收集。
原摘要

Collecting whole-body demonstrations for humanoid manipulation mostly relies on teleoperation, which is costly and hard to scale up. The Universal Manipulation Interface (UMI) provides a scalable data collection paradigm, but end-effector trajectories alone underdetermine humanoid whole-body coordination, which is insufficient for whole-body demonstration collection. Therefore, we introduce Whole-Body UMI (WB-UMI), a task-agnostic, real-time and end-effector conditioned motion generator that decouples whole-body coordination learning from task semantics learning through a shared end-effector interface. A diffusion policy learns from native UMI demonstrations, while WB-UMI learns independently from retargeted motion capture, requiring no body trackers or paired image--whole-body demonstrations during task-specific data collection. In real deployment, an asynchronous hierarchy integrates the diffusion policy, motion generator, and a whole-body controller with latency compensation and measured-state feedback. Real-robot experiments on G1 support real-time closed-loop transfer across four tasks, achieving 90% success in drawer closing, 80% in shelf pick-and-place, 30% in ball toss, and 40% in Loco-PnP, which shows the effectiveness of this hierarchy in transferring native UMI skills to humanoid whole-body manipulation.

Lucky Kant Nayak, Narayanan Palghat Parameswaran, Neehar Peri, Deva Ramanan
总结: 提示到轨迹生成框架,学习四足动态技能。
方法: 编码代理从技能提示生成参考轨迹,指导示例强化学习。
证据: 比奖励塑形更易泛化多样技能与形态。
为什么适合我: 扩展四足敏捷运动模仿与参考动作跟踪。
原摘要

We present MimicAgent, a prompt-to-trajectory generation framework for learning dynamic quadruped skills. Although reward shaping is extensively used when training quadruped policies, navigating the resulting reward landscape is notoriously difficult, requiring hours of "graduate student descent". Eureka attempts to automate reward design with LLMs, but we find that it struggles to generalize across diverse skills and morphologies. Our key observation is that it is far easier for a human - and by association, an LLM - to generate reference motions than to shape reward functions. Our hypothesis is motivated by the success of example-guided RL for humanoids, which exploits large-scale motion capture datasets as references for training locomotion policies. Unlike humanoids, quadrupeds lack such reference motion data. Towards this end, we propose MimicAgent, an agentic harness that, given a skill prompt, generates quadruped reference trajectories with coding agents. These coarse reference trajectories are then used to train example-guided RL policies that are deployable in simulation and in the real-world. Notably, we find that when prompting Claude Fable 5.1 within our agentic harness, 87% of prompts yield semantically aligned reference trajectories.

Haruto Nagahisa, Kohei Matsumoto, Yuki Hyodo, Ryo Kurazume
总结: 扩散转向实现社交导航中性能保持的在线适应。
方法: 固定扩散策略仅训练噪声策略,整合多种子基策略。
证据: 相比基线更好保持导航性能并适应部署环境。
为什么适合我: 扩散策略在线适应利于真实机器人迁移。
原摘要

In social navigation, modeling the complex interactions between humans and robots is difficult, and deep reinforcement learning has therefore been actively studied. However, because simulation alone cannot fully reproduce diverse scenarios, robot dynamics, and the social conventions that vary across deployment environments, fine-tuning in the deployment environment is promising. In doing so, learning that preserves the base model's performance is required, so as not to compromise the primary objective of navigation, namely avoiding pedestrians and reaching the destination. In this study, we propose a method that applies diffusion steering via reinforcement learning (DSRL), which trains only the noise policy while keeping the diffusion policy fixed, thereby achieving learning that preserves performance. Furthermore, we integrate diffusion-based RL policies trained with multiple seeds to construct the base policy, improving learning performance. Our evaluation shows that, compared with other methods, the proposed method enables efficient learning while preserving performance, and we confirm flexible behavior control through adaptation to social conventions, as well as its effectiveness on a physical robot through hardware-in-the-loop simulation.

Lalit Jayanti, Kashu Yamazaki, Yuto Shibata, Kotaro Amaya, Katerina Fragkiadaki
总结: 噪声空间轨迹优化生成可扩展人形交互运动参考。
方法: 优化预训练文本运动模型噪声,满足接触与场景约束。
证据: 生成运动可被跟踪执行并训练深度条件视觉策略。
为什么适合我: 支持接触丰富交互参考生成与运动模仿。
原摘要

Humanoid robots can acquire complex skills by imitating kinematic humanoid motion references, yet reliable references for contact-rich interactions remain difficult to obtain: motion capture deteriorates under occlusion and close physical contact, while retargeting introduces additional contact and geometric inconsistencies. We present HIGenNTO, a framework that synthesizes humanoid-scene interaction motion references by optimizing the initial noise of a pretrained text-conditioned motion model under sparse spatiotemporal and scene constraints. The same formulation satisfies desired contacts, avoids collisions, and maintains stable support while retaining the prior's realism and temporal coherence, generating interaction motions from scratch and composing long-horizon behaviors stage-wise. Across robot-environment and robot-object tasks, HIGenNTO produces motions that can be executed by tracking policies in simulation and used to train depth-conditioned visuomotor policies operating solely from onboard sensing. We deploy these policies on a Unitree G1 across four contact-rich tasks. Finally, the task specifications themselves can be written by a coding agent, which proposes interaction tasks and compiles them into prompt, constraint, and scene programs, authoring three of our eight evaluated tasks and four further behaviors. Together, these results establish a scalable path from high-level task descriptions to physically executable humanoid interactions.

Lei Ye, Haibo Gao, Yitang Li, ... , Hao Zhao, Liang Ding
总结: 预测动作扩散策略,实现可引导的人形机载控制。
方法: 基于本体感知联合生成可执行动作与未来状态轨迹。
证据: 支持测试时运动引导,无需特权全状态。
为什么适合我: 利于生成式策略跟踪多样参考并全身控制。
原摘要

Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.