Papers for 2026-10-06

10 papers
Ziqi Han, Yitang Li, Junhan Sun, ... , Xue Wang, Hao Zhao
总结: 提出首个人形物体交互行为基础模型I-BFM,单一策略经潜变量条件化即可执行搬运、推、踢等闭环交互,无需任务专属优化。
方法: 用前向-后向表征与无监督强化学习学习人形、物体与接触的耦合动力学共享表征,并以短时交互目标与长时目标双尺度训练。
证据: 单一策略直接以奖励条件化完成搬运、推踢、目标到达与运动跟踪等多种下游交互任务,验证了潜命令驱动的闭环控制泛化能力。
为什么适合我: 为我提供了超越参考动作跟踪的生成式全身控制路径,奖励条件化可让接触丰富操作的控制器免于逐任务重训。
推荐理由: 直接对应人形全身人-物交互:用无监督强化学习学耦合接触动力学,并以奖励条件化同一策略完成闭环交互,贴近全身控制与技能表征。
原摘要

Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.

Yucheng Zhang, Sirui Xu, Jinhong Li, ... , Yu-Xiong Wang, Liang-Yan Gui
总结: 自演化运动模仿框架,让机器人参考数据与跟踪策略互相改进,大规模扩展带灵巧手的人形全身移动操作。
方法: 将人-物体交互动捕数据重定向为带灵巧手的机器人参考,训练物理仿真通用跟踪器,再任务保持式微调参考形成数据飞轮。
证据: 构建了规模与多样性超越以往人形跟踪系统的移动操作参考库,通用跟踪器可在仿真中执行该大规模数据集。
为什么适合我: 运动重定向加数据飞轮正好契合我用大规模参考动作学全身控制器的目标,其自演化机制可低成本扩充参考库。
推荐理由: 高度契合人形移动操作:将人-物交互重定向为全身参考,并用物理跟踪策略与数据飞轮扩展灵巧loco-manipulation。
原摘要

Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.

Seungho Yeom, Zhenyu Wu, Jaeyoung Huh, ... , Soofiyan Atar, Michael Yip
总结: OCLO无需人类运动数据的人形移动操作系统,仅凭两个末端目标指令即可在线生成全身姿态并实现柔顺协调控制。
方法: 以解析可达性先验在线生成骨盆高度与躯干朝向,经策略在环采样细化,并用弹簧阻尼模型位移末端参考学习全身柔顺性。
证据: 仿真中可达性先验使命令达成成功率从37.8%升至77.8%,且腿部、腰部与骨盆能对外部负载做出顺从性让步。
为什么适合我: 证明摆脱动捕依赖也能学到柔顺移动操作,其在线姿态生成与力柔顺机制可增强我在非结构化环境的操作鲁棒性。
推荐理由: 无人类动作数据的人形loco-manipulation,在线生成全身姿态并学习力顺应,贴近全身协调与接触交互。
原摘要

Most humanoid loco-manipulation controllers require human motion data to learn whole-body coordination and posture, leaving policies reliant on external sources to provide this data. We present OCLO (Online-posture Compliant LOco-manipulation), a humanoid loco-manipulation system trained without human motion data and commanded only through two end-effector targets. Because these targets do not uniquely determine whole-body posture, OCLO generates pelvis height and torso orientation online using an analytic reachability prior, further refined through policy-in-the-loop sampling with a task-agnostic cost. OCLO also learns whole-body compliance by displacing end-effector references according to measured forces through a spring-damper model, encouraging the legs, waist, and pelvis to yield to external loads. In simulation, using the reachability prior leads to a 77.8% success rate in acquiring the commanded reference, a vast improvement over the 37.8% success rate accomplished without the prior. Further, refinement reduces end-effector orientation error across all evaluated tasks. The same posture module improves a pretrained SONIC controller on four of five tasks. Without compliance training, policies tend to lose balance under disturbances rather than sacrifice tracking. On a Unitree G1, OCLO maintains balance under end-effector disturbances that cause its ablations to fail and performs seven loco-manipulation tasks, including crouched walking and picking up a box from a low surface. Project website: https://oclo-humanoid.github.io/

Continual Humanoid Motion Learning

5.0/5 很相关 裁判分 4.6 Candidate 当日相对
Zhewen He, Hao Huang, Geeta Chandra Raju Bethala, ... , Anthony Tzes, Yi Fang
总结: 研究人形全身运动的持续学习,单一控制器从顺序任务流学新技能而不遗忘旧技能,且无需回放历史数据。
方法: 提出相似度引导的LoRA渐进神经网络,用动态时间规整加最优传输的二级动作相似度决定继承哪个旧技能及分配多少新容量。
证据: 在六个顺序学习技能类别上取得最佳前向迁移(0.125对0.079),凭结构设计防止灾难性遗忘并大幅提升训练效率。
为什么适合我: 为我解决全身控制器技能持续扩展问题,轻量LoRA模块可让参考动作库随时间增长而不侵蚀已掌握技能。
推荐理由: 直接研究人形全身运动控制器的持续技能学习与抗遗忘,契合通用运动跟踪和技能复用。
原摘要

Humanoid whole-body controllers can now track a diverse set of dynamic motions, but they are typically trained offline and then frozen, so teaching such a controller a new skill tends to erode the skills it already mastered. We study continual learning for humanoid whole-body motion, where a single controller must acquire skills from a sequential task stream without revisiting past data. We introduce Similarity-guided LoRA-PNN, a progressive neural network (PNN) policy that prevents catastrophic forgetting by construction while reusing knowledge across skills through lightweight low-rank adaptation. A two-level motion-similarity measure, built from dynamic time warping aggregated by optimal transport, decides which prior skill to build on and how much new capacity to allocate, yielding strong forward transfer and large efficiency gains. Across six sequentially learned skill categories, our similarity-guided LoRA policy attains the best forward transfer (0.125 vs. 0.079) and the highest average accuracy among all methods, while saving up to 94.5% of trainable parameters and 40.8% of training time. The resulting controller reaches 96.13% sim-to-sim transfer and is deployed on a physical Unitree G1. Our code is available at https://anonymous.4open.science/r/continual-humanoid-learning-35D3.

Echo in the Steps: Learning Perceptive Humanoid Parkour with Gated Memory

5.0/5 很相关 裁判分 4.4 Candidate 当日相对
Ming-Ju Lee, Zizhuo Wang, Shaoting Zhu, ... , Hang Zhao, Yiming Li
总结: 基于机载深度感知的人形跑酷框架,在稀疏落脚点与高度不连续地形上实现稳定而敏捷的快速穿越。
方法: 显著性引导的时序感知模块结合门控记忆跨帧保留深度特征,从部分观测实现可靠落足,并用交替损失正则化步态对称。
证据: 大量实验显示在落脚点受限地形上的穿越能力显著优于基线,仅用机载深度观测即可完成敏捷跑酷任务。
为什么适合我: 直接对应我的跑酷与感知运动方向,门控记忆处理部分可观测下落足选择的思路可用于我的地形利用控制器。
推荐理由: 完全命中人形感知跑酷:机载深度、落足选择与稀疏支撑地形上的敏捷全身运动。
原摘要

While recent advances in perceptive locomotion have enabled humanoid robots to traverse structured terrains, agile parkour in highly discontinuous environments remains an open challenge. In particular, crossing sparse footholds and narrow support regions requires precise foothold selection, effective use of visual observations, and consistent alternating foot placement during fast transitions. In this paper, we present a perceptive humanoid parkour framework that enables stable traversal across terrains with limited foothold availability using only onboard depth observations. The framework features a saliency-guided temporal perception module that combines a saliency prior with gated memory. It retains informative depth features across frames, enabling reliable foot placement from partial observations. By introducing an alternation loss, our symmetry regularization encourages alternating gait patterns and improves traversal robustness. Extensive experiments show that our method significantly improves success rate and foothold accuracy on challenging terrains in both simulation and the real world.

Humanoid Rickshaw Pulling: Whole-Body Locomotion under Coupled Wheeled Loads

5.0/5 很相关 裁判分 4.3 Candidate 当日相对
Yangzhi Yang, Xiansheng Lin, Zhaoming Xie, Xiaobin Xiong
总结: 人形拉黄包车式重载全身控制,在未知耦合轮式负载动力学下跟踪车辆运动,同时保持平衡与稳定抓握。
方法: 特权教师利用车辆状态、交互力与负载属性训练,其动作与潜变量被蒸馏入基于历史的本体感知学生,再强化学习微调。
证据: 与无历史和仅历史基线相比,策略实现精确车辆跟踪并降低速度误差,能适应构型相关的未知耦合负载力。
为什么适合我: 拉重载是我移动操作的重要延伸,历史条件化隐式推断耦合动力学的方法可迁移到我处理外部负载的接触任务。
推荐理由: 人形在持续上肢接触和未知轮式载荷下保持平衡与抓持,直接对应接触丰富的全身运动与移动负载交互。
原摘要

Humanoid robots could transport payloads substantially heavier than themselves by pulling passive wheeled vehicles instead of carrying the load. This capability, however, creates a coupled locomotion problem: the robot must maintain persistent upper-body contact while adapting to unknown, configuration-dependent forces arising from the payload, vehicle, and terrain. We present a whole-body control framework for humanoid rickshaw pulling that tracks commanded vehicle motion while preserving balance and stable grasps under uncertain load dynamics. During training, a privileged teacher exploits vehicle states, interaction forces, and load properties. Its actions and latent are distilled into a history-conditioned student that implicitly infers coupled dynamics from proprioceptive responses, followed by reinforcement-learning fine-tuning. Comparisons with \emph{No History} and \emph{Only History} baselines show that the resulting policy achieves accurate vehicle tracking while reducing vehicle oscillation, torso tilt, and actuation cost. Behavioral analysis shows that Unitree G1 propels the rickshaw and generates gait-synchronized whole-body reactions that stabilize its lateral and roll motions. Moreover, pulling redistributes joint effort and yields a lower robot-normalized cost-of-transport proxy than unloaded walking over most tested load--speed conditions. On hardware, a single policy performs starting, sustained pulling, turning, and stopping with both rigid payloads and human passengers, handling a loaded rickshaw mass of up to 115~kg without load-specific retuning. These results demonstrate robust heavy-load transportation through coordinated and persistent humanoid--vehicle interaction.

Jin Chen, Yiming Jiang, Chongyang Xu, ... , Steven Hoi, Hongyang Li
总结: 首个自我中心人-人形技能迁移框架,将人类演示中的协调全身移动操作技能零样本迁移到真实机器人VLA策略。
方法: 粗到细动作对齐结合运动学参考校正与动力学感知细化,并用机器臂渲染与训练时图像增强缩小视觉具身差异。
证据: 四个真实任务上,对齐人类数据训练的VLA策略实现零样本技能迁移,任务得分与遥操作机器人数据训练的策略相当。
为什么适合我: 提供免遥操作规模化获取全身操作数据的路径,其动作对齐方法可直接改进我的运动重定向与模仿管线。
推荐理由: 从第一人称人类演示迁移协调的全身loco-manipulation技能,核心是跨本体动作对齐与重定向。
原摘要

Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.

CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video

5.0/5 很相关 裁判分 4.1 Candidate 当日相对
Zhuoqun Chen, Shucheng Jia, Boyuan Chen
总结: 从视频学习反应式柔顺人-人形交互,以双人舞为载体让人形与移动伙伴协调步法并维持双手持续接触。
方法: 将双人舞视频重定向为机器人参考与移动伙伴,提出多链路柔顺增强,在双手结构化外力下调整参考生成力感知训练数据。
证据: 仿真中策略能跟随观测到的伙伴运动、保持演示的移动风格,并对物理交互做出柔顺响应。
为什么适合我: 视频到力感知数据的柔顺增强思路,可为我的全身运动跟踪注入接触力响应,提升人机交互的自然与稳定。
推荐理由: 从视频重定向双人舞蹈并学习力顺应的人-人形持续接触,契合运动重定向与反应式全身交互。
原摘要

Partnered human-humanoid interaction couples locomotion with continuous physical contact. A humanoid needs to coordinate with a person's motion while responding to interaction forces and maintaining stable and natural movement. We present CoDance, a framework for learning reactive and compliant human-humanoid interaction from video. We study partnered dancing as a challenging instantiation, where a humanoid coordinates its footsteps with a moving partner and maintains continuous two-hand contact. Given a single video of two human dancers, CoDance retargets their motions into a robot reference and a moving partner. We introduce a multi-link compliance augmentation that transforms the kinematic demonstration into force-aware training data by adapting the robot reference under structured forces at both hands. Policies trained on this data follow the observed partner while preserving the demonstrated locomotion style and responding compliantly to physical interaction. In simulation, the policies adapt their footsteps to changes in the partner and reproduce approximately 80% of the wrist displacement encoded by the augmented demonstrations. On a physical humanoid, CoDance enables sustained two-hand dancing with a human partner including repeated transitions between forward and backward motions.

Nobuo Namura, Masayuki Hiromoto, Kento Uemura, Hironobu Sasaki, Kanata Suzuki
总结: 四足机器人运输无固定堆叠货箱,用多目标强化学习在线权衡运动敏捷性与负载稳定性,无需专用负载传感器。
方法: 训练以偏好向量条件化的多目标基础策略,再在冻结策略上训练权重调节器,从本体感知在线调整运动与稳定偏好。
证据: 仿真中整体运输成功率与单目标基线相当或更优,包括未见过的三箱堆叠,尽管仅用两箱配置训练。
为什么适合我: 偏好条件化加在线权重调整可解决我敏捷运动与负载稳定的冲突,对四足与足式移动操作任务均具参考价值。
推荐理由: 属于四足强化学习运动控制,关注不平地形上未固定载荷的稳定运输;相关但偏载荷偏好权衡,而非通用全身跟踪或跑酷。
原摘要

Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.

Sebastian Hirt, Lukas Theiner, Jan Peters, Rolf Findeisen
总结: 面向分层控制系统的上下文贝叶斯优化框架,联合调优规划、全身运动与底层控制各层分布参数。
方法: 保留任务性能、实现质量与控制量的分离观测,用相关多输出高斯过程建模,并解析评估其向闭环总目标的聚合。
证据: 在人形分层控制器调优中捕获跨层参数交互,相比独立调参或标量黑盒建模获得更好的闭环整体性能。
为什么适合我: 我的全身控制器多为分层结构,该上下文贝叶斯优化可高效联合调优各层参数,避免忽略跨层耦合效应。
推荐理由: 任务落在人形loco-manipulation的分层全身控制,但是用上下文贝叶斯优化联合调参,而非学习可跟踪多样动作的策略。
原摘要

Hierarchical control architectures are widely used to decompose complex control problems into interacting control levels and are particularly important in robotics, where planning, whole-body motion, and lower-level control must be coordinated across different levels of abstraction and time scales. Their overall closed-loop performance, however, depends strongly on parameters distributed across the hierarchy, such that tuning controllers on different levels independently may neglect relevant cross-layer interactions. We propose a contextual Bayesian optimization framework for joint parameter learning in hierarchical control systems. Rather than modeling closed-loop performance only as a scalar black-box function, we retain separate observations of task performance, realization quality, and control effort. A correlated multi-output Gaussian process models these performance components, while their known aggregation into the overall closed-loop objective is evaluated analytically. The formulation exploits three complementary consequences of hierarchical control: informative performance quantities exposed by the hierarchy, coupling between parameters of different controller levels, and variations of these relations with operating conditions. We evaluate the approach for humanoid loco-manipulation, jointly tuning a centroidal predictive controller and a whole-body controller for physical box pushing under varying box mass. The proposed method achieves the lowest mean empirical regret during both training and adaptation among the considered baselines.