Papers for 2026-10-01

10 papers
Guillaume Besset, Erwann Carn, Timothée Carecchio, ... , Justin Carpentier, Ajay Suresha Sathya
总结: 提出统一框架将人类与物体交互动作联合重定向到人形机器人,解决移动操作中接触保持与形体差异难题。
方法: 以符号距离、最近表面点与相对方向表示表面交互,用熵最优传输在人、机器人、物体几何间传递交互量,并纳入约束逆运动学求解。
证据: 无需固定物体轨迹即可在形体差异下保持地面与物体接触一致,约束IK在接触保持与运动跟踪间取得平衡。
为什么适合我: 运动重定向是我构建全身控制器的核心环节,该方法为接触丰富的人形移动操作示教迁移提供了直接可用的工具。
原摘要

Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.

Hyeonjin Choi, Joongheon Kim, Daekyum Kim
总结: 受被动动态行走启发,训练早期构造斜坡等效力场辅助发现节能步态,随后移除指导再做标称动力学优化。
方法: 早期用倾斜重力场配合课程耦合奖励引导矢状推进,无需参考轨迹、步态相位或接触时刻表,后期去除全部PDW专用引导。
证据: 在29自由度Unitree G1五种子研究中,0.5–2.0m/s指令速度下机械运输成本降低6.8–15.2%,速度跟踪不退化。
为什么适合我: 无参考的节能双足步态可直接提升真实人形续航与步态自然性,为我的敏捷行走控制器能耗设计提供新思路。
原摘要

Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagittal progression on flat collision geometry, complemented by curriculum-coupled reward terms. The core framework requires no reference trajectories, gait phases, or contact schedules. In a five-seed forward-locomotion study on a 29-DoF Unitree G1, the framework reduces mechanical cost of transport by 6.8-15.2% over commanded speeds of 0.5-2.0m/s without degrading velocity tracking. Mechanical-work decomposition attributes the reduction to positive actuator work, and reward-matched comparisons separate the guided regime's faster gait acquisition from the tilt's additional benefit to converged economy. The framework extends to unassisted omnidirectional locomotion, where its benefit persists once a walking-specific motion prior supplies kinematic coordination, the combination reducing speed-matched cost of transport by 18.7%. On hardware, forward cost of transport falls by 16.3% with the motion prior and by 4.5% without it, the latter within the trial-to-trial spread.

Yiming Jiang, Chen Jin, Chongyang Xu, ... , Aimin Hao, Yisheng He
总结: 提出EgoAlign,把自我中心人类示教转化为通用连续全身控制器可用的动作与状态监督,用于长程人形移动操作。
方法: 用目标机器人模型与仿真器以执行反馈引导示教采集,尺度对齐细化上身交互几何,因果回放重建机器人状态与运动token标签。
证据: 仅凭适配后的人类示教微调视觉-语言-动作模型即可支撑长程移动操作,全程无需采集物理机器人示教。
为什么适合我: 契合我对感知运动与遥操作的兴趣,提供了从人类视频低成本扩展人形全身控制监督数据的可扩展管线。
原摘要

Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion references for visually guided periodic stepping while adapting upper-body interaction geometry through scale alignment and controller-in-the-loop refinement. A final causal replay reconstructs the corresponding robot states and motion-token labels for training with the human observations. We assess the resulting supervision by fine-tuning a vision--language--action model solely on adapted human demonstrations and deploying it zero-shot on a physical humanoid. The resulting policies perform long-range object relocation, navigation to unseen goal positions, and independently evaluated foot interaction. Refinement improves simulated hand alignment and physical pickup success over kinematic alignment alone, while human collection reduces on-site acquisition time relative to teleoperation. https://lambdahumanoid.github.io/EgoAlign/

Tan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger, ... , Xinchao Wang, An T. Le
总结: 将潜行为空间视为跨形体可迁移资产,一次蒸馏即得到多人形共享、可提示的行为基础模型空间。
方法: 利用重定向提供的帧级跨形体对应,采用无机器人特定参数的统一编码器同时蒸馏所有形体,潜变量条件跟踪器输出全身控制。
证据: 训练耗时从数百GPU小时降到不足1小时,单一向量即可表示待模仿动作、目标姿态或待最大化奖励。
为什么适合我: 运动模仿与全身跟踪是我的主线,跨形体共享潜空间让行为先验能在不同人形间复用迁移,极具启发。
原摘要

Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the transferable asset for various embodiments. As retargeting provides frame-level cross-embodiment correspondence, we propose a unified encoder architecture with no robot-specific parameters for distilling the behavior space to address all training embodiments simultaneously in less than a GPU-hour. Following this encoder, latent-conditioned trackers turn the distilled latent into whole-body control in a conventional PPO training manner in just 10 more GPU-hours. On three distilled humanoids, all three prompting modes transfer: motion tracking with latent-conditioned policy losing only $0.025$ rad to its joint-conditioned counterpart, smooth goal reaching between poses with no falls, and reward optimization for all $41$ reward prompts. Our experiments further reveal that 1) regressing the encoder on a quarter of the motion corpus costs only $5\%$ of tracking performance and 2) training the encoder on a subset of robots and evaluating on an unseen one recovers up to $89\%$ of the tracking performance of seen robots, demonstrating cross-embodiment generalization to morphologically similar robots. We also verify the pipeline on real robots across all three prompting modes and with flow-based generated latents. Project website: https://dotandung.github.io/crossbfm/

Ivan Ovinnikov, Pascal Sutter, Christian Gehring, Jordis Herrmann
总结: 提出预测安全课程,用学到的未来安全代价预测分配足式训练经验,减少平均性能高却仍有罕见灾难性失效的策略。
方法: 从策略回滚训练分布式安全评论家,按其预测优先采样地形上下文与历史随机化事件,仅改训练分布而不动奖励与优化损失。
证据: 在受控崎岖地形与生产级行走系统上均优于标准地形进阶、优势回放与学习进度课程,难地形与退化条件下收益最大。
为什么适合我: 把控制器可靠部署到非结构化地形是我的核心目标,该课程机制可与现有RL训练直接叠加以提升鲁棒性与安全性。
原摘要

Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part because standard curricula primarily adapt task difficulty rather than the distribution of safety-critical experience. We introduce Predictive Safety Curricula (PSC), a framework for allocating locomotion training experience using learned predictions of future safety cost. PSC trains a distributional safety critic from policy rollouts and uses its predictions to prioritize both terrain contexts and previously encountered randomized events. The resulting curriculum modifies the training distribution while leaving the task reward and policy-optimization loss unchanged. We evaluate PSC in controlled rough-terrain locomotion and in production locomotion systems. PSC improves reliability relative to standard terrain progression, advantage-based replay, and learning-progress curricula, with the largest gains on difficult terrain and under degraded observations. The same allocation principle transfers to two production locomotion stacks. On ANYmal-D hardware, PSC reduces shank-collision incidence by $63\%$ relative to the learning-progress curriculum across three matched training seeds, with a reduction in every seed. On a production stair-climbing platform, PSC eliminates observed shank collisions in the evaluated hardware trials. These results show that learned predictions of future safety cost can provide an effective signal for allocating training experience toward rare failure modes and improving locomotion reliability.

Nico Bohlinger, Jan Peters
总结: 给评论家嵌入显式时间几何,使嵌入距离以到达目标所需时间为单位,解决空间近但时间远的目标度量问题。
方法: 基于生存强化学习,把未达目标与其他轨迹的目标嵌入推远至少一个折扣视野,并预测到达时间的完整分布与在目标附近停留时长。
证据: 时间几何使表示空间中的距离具备时间语义,并显式建模快速到达一次不等于可靠到达这一直觉。
为什么适合我: 跑酷与崎岖地形中到达代价由地形与自身能力决定,时间感知评论家可改进我的地形利用与技能选择策略。
原摘要

A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so the policy should not follow the temporal distance directly. Instead, we build on survival reinforcement learning and predict from our temporal embeddings not only the full distribution of goal-reaching times but also the time spent near the goal. Thereby, the policy is trained to favor actions that reach the goal sooner and more reliably and that keep the agent near it. ChronoSRL learns faster and reaches higher performance than contrastive, action-chunked contrastive, and survival reinforcement learning baselines on seven standard locomotion and navigation benchmarks, even with much smaller networks. To test the limits of self-supervised reinforcement learning, we introduce velocity tracking, goal-position reaching, and box climbing tasks with a quadruped robot in a realistic sim-to-real locomotion setup, and show how the shaping terms that are typical for robotics can be naturally incorporated into our framework. ChronoSRL is the only one of the tested self-supervised reinforcement learning methods that learns to stay at the commanded velocities and goal positions, and climbs the highest boxes.

Abu Hanif Muhammad Syarubany, Chang D. Yoo
总结: 将SIM(3)等变点云编码器引入DP3,为43关节Unitree G1构建少示教即可泛化的人形移动操作视觉运动策略。
方法: SIM(3)等变VNN编码的扩散规划器以6Hz输出全身指令块,由冻结预训练RL行走策略与差分IK手臂模块以50Hz执行,行为克隆端到端训练。
证据: 在两个IsaacLab仿真基准上,5–100条示教下胜过四个非等变基线,泛化到不同物体位姿与光照。
为什么适合我: 高层扩散规划加冻结RL底盘的分层接口正是我关注的移动操作范式,等变编码显著提升数据效率,值得借鉴。
原摘要

Visuomotor policies for humanoid loco-manipulation must generalize across object poses and lighting from only a handful of demonstrations. 3D Diffusion Policy (DP3) conditions a diffusion-based action generator on point-cloud features, but its PointNet-style encoder has no built-in equivariance to the rotations, translations, and scalings (SIM(3)) that manipulation tasks respect. EquiBot closed this gap for wheeled manipulators with a SIM(3)-equivariant Vector Neuron Network (VNN) encoder. We extend this to a substantially more complex embodiment, the 43-joint Unitree G1 humanoid, and propose EquivDP3: a two-stage policy where a high-level diffusion planner with a SIM(3)-equivariant VNN encoder emits 6 Hz whole-body command chunks, executed at 50 Hz by a frozen, pre-trained RL locomotion policy and a differential inverse-kinematics module for the arms, trained end-to-end by behavior cloning. Across two simulated IsaacLab benchmarks and four non-equivariant baselines (5-100 demonstrations, in- and out-of-distribution), EquivDP3's advantage concentrates in the low-data regime: at 5-10 demonstrations it reaches 67.1% success versus 38.2-52.4% for the baselines, while by 50-100 all encoders converge (74.3-85.2%) and the ordering is no longer meaningful. A proprioception-only control confirms this gap is genuinely perceptual: with the point cloud removed, success drops to 31% vs. 60% (EquivDP3) at 5 demonstrations and 78% vs. 99% at 10, but vanishes by 50-100, showing the high-data plateau reflects a benchmark ceiling, not five encoders learning the same invariance. The encoder costs only 0.8 ms of extra latency per action chunk over the PointNet encoder it replaces. Baking geometric symmetry into a hierarchical diffusion policy's perception backbone is a practical, nearly free way to improve data efficiency for humanoid loco-manipulation when demonstrations are scarce.

Feiyang Wu, Chenxiao Gao, Chen Yang, ... , Bo Dai, Anqi Wu
总结: 提出谱技能作为规划器与控制器间的潜在命令接口,以预测式学习紧凑编码短运动段,并支持技能组合。
方法: 通过预测后续运动而非重建输入学习运动表征,在29自由度人形上训练以谱技能为条件的全身跟踪控制器。
证据: 全局跟踪误差较SOTA降低62%;同一冻结控制器无需过渡策略即可链接技能,并通过正交方向相加组合出新行为。
为什么适合我: 可组合、易预测的运动表征正是构建可跟踪多样参考动作的全身控制器所需接口,直接服务我的技能复用目标。
原摘要

Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideally allow new behaviors to be composed from prior ones. In this work, we introduce spectral skills, a latent representation of this interface that meets these requirements through predictive representation learning. By design, spectral skills compactly encode short motion segments and are learned by predicting subsequent motion rather than reconstructing the encoder input. On a 29-DoF humanoid, a controller conditioned on spectral skills reduces global tracking error by 62\% relative to the state of the art. The same frozen controller chains independently encoded skills without a separate transition policy. It also composes new behaviors by adding orthogonal directions to any compatible base skill, producing combinations unseen in the training data. We demonstrate tracking, chaining, and composition, as well as control through a language-conditioned planner, on Unitree G1 hardware. Project page: https://spectral-skill.github.io

Zihan Wang, Zhen Wu, Pieter Abbeel, ... , Guanya Shi, Angjoo Kanazawa
总结: 提出PRISM,用视频生成从少量真实视频扩增大量反事实人-物交互视频,训练可泛化的人形移动操作策略。
方法: 真实-仿真-真实管线:V2V生成数百条多样化反事实视频,再以接触锚定方式重建人与物体运动并重定向为物理可行轨迹。
证据: 反事实视频的类内多样性使单一策略即可泛化到各类别中未见物体,验证了从极少真实视频扩展训练数据的路径。
为什么适合我: 生成式数据扩增结合运动重定向,与我用生成策略扩展训练数据的思路一致,可缓解接触丰富交互示教的采集瓶颈。
原摘要

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

Takumi Asada, Hideo Furuhashi, Kenta Tabata, Renato Miyagusuku, Koichi Ozaki
总结: 面向多关节水下仿生机器人,提出基于传感模态的多模态控制机制,在同一结构上实现游动与步态能力。
方法: 采用四条各具四轴腿鳍的机构,控制器不依赖预定义条件规则,而是依据传感器模态生成非线性行为选择。
证据: 基于多传感器数据与多种行为模式完成物理与仿真双重测试,展示了超越预设行为范围的能力表达。
为什么适合我: 传感器驱动、无需规则库的行为生成对腿式控制有一定借鉴意义,但对象偏离人形与陆地全身控制主线,价值有限。
原摘要

Multimodal biomimetic underwater robots (BURs) can conduct underwater tasks suitable for the environment. Combining the characteristics of aquatic organisms enables the swimming and leggedlocomotion required for underwater exploration. Locomotion control mechanism relies on rule-based behavior selection and the designer's discretion. This limits the robot's ability to acquire new behavioral capabilities to the predetermined range of behaviors. To address these challenges, we propose a mechanism and control system that enables the expression of multimodal locomotion capabilities from the same multi-jointed structure. A mechanism equipped with four leg-fins each having four axes is used. This controller achieves nonlinear behavior based on sensor modalities, rather than relying on predefined conditional rule-based on locomotion functions. This system was validated through both physical and simulation testing based on multiple sensor data and behavioral patterns. Utilizing a potential function in multimodal locomotion control was verified to enable transitions between two or three behaviors. Implementing the control method as a multimodal controller is expected to enhance its application in underwater exploration. Our project page is at https://tasada038.github.io/multi-jointed-bur/.