Papers for 2026-02-06

10 papers
Zhanxiang Cao, Liyun Yan, Yang Zhang, ... , Cewu Lu, Yue Gao
总结: HiWET通过将人形机器人移动操作重构为世界坐标系下的末端执行器跟踪,并结合分层强化学习与运动学流形先验,实现了长时程任务中更精确、稳定且可迁移到真实机器人的移动操作控制。
原摘要

Humanoid loco-manipulation requires executing precise manipulation tasks while maintaining dynamic stability amid base motion and impacts. Existing approaches typically formulate commands in body-centric frames, fail to inherently correct cumulative world-frame drift induced by legged locomotion. We reformulate the problem as world-frame end-effector tracking and propose HiWET, a hierarchical reinforcement learning framework that decouples global reasoning from dynamic execution. The high-level policy generates subgoals that jointly optimize end-effector accuracy and base positioning in the world frame, while the low-level policy executes these commands under stability constraints. We introduce a Kinematic Manifold Prior (KMP) that embeds the manipulation manifold into the action space via residual learning, reducing exploration dimensionality and mitigating kinematically invalid behaviors. Extensive simulation and ablation studies demonstrate that HiWET achieves precise and stable end-effector tracking in long-horizon world-frame tasks. We validate zero-shot sim-to-real transfer of the low-level policy on a physical humanoid, demonstrating stable locomotion under diverse manipulation commands. These results indicate that explicit world-frame reasoning combined with hierarchical control provides an effective and scalable solution for long-horizon humanoid loco-manipulation.

Weidong Huang, Jingwen Zhang, Jiongye Li, ... , Yaodong Yang, Yao Su
总结: ECO将能耗与参考运动从奖励中分离并作为显式约束进行强化学习优化,使人形机器人在保持稳定鲁棒行走的同时显著降低能耗,并在BRUCE机器人仿真与实机实验中优于MPC、奖励塑形RL和多种约束RL基线。
原摘要

Achieving stable and energy-efficient locomotion is essential for humanoid robots to operate continuously in real-world applications. Existing MPC and RL approaches often rely on energy-related metrics embedded within a multi-objective optimization framework, which require extensive hyperparameter tuning and often result in suboptimal policies. To address these challenges, we propose ECO (Energy-Constrained Optimization), a constrained RL framework that separates energy-related metrics from rewards, reformulating them as explicit inequality constraints. This method provides a clear and interpretable physical representation of energy costs, enabling more efficient and intuitive hyperparameter tuning for improved energy efficiency. ECO introduces dedicated constraints for energy consumption and reference motion, enforced by the Lagrangian method, to achieve stable, symmetric, and energy-efficient walking for humanoid robots. We evaluated ECO against MPC, standard RL with reward shaping, and four state-of-the-art constrained RL methods. Experiments, including sim-to-sim and sim-to-real transfers on the kid-sized humanoid robot BRUCE, demonstrate that ECO significantly reduces energy consumption compared to baselines while maintaining robust walking performance. These results highlight a substantial advancement in energy-efficient humanoid locomotion. All experimental demonstrations can be found on the project website: https://sites.google.com/view/eco-humanoid.

Wandong Sun, Yongbo Su, Leoric Huang, ... , Ethan Xie, Zongwu Xie
总结: 本文提出一种从原始深度像素端到端学习人形机器人视觉行走的框架,通过高保真深度传感器仿真、视觉感知行为蒸馏和按地形划分的多奖励/多判别器训练,实现了在真实机器人上跨复杂地形的鲁棒运动。
原摘要

Achieving robust vision-based humanoid locomotion remains challenging due to two fundamental issues: the sim-to-real gap introduces significant perception noise that degrades performance on fine-grained tasks, and training a unified policy across diverse terrains is hindered by conflicting learning objectives. To address these challenges, we present an end-to-end framework for vision-driven humanoid locomotion. For robust sim-to-real transfer, we develop a high-fidelity depth sensor simulation that captures stereo matching artifacts and calibration uncertainties inherent in real-world sensing. We further propose a vision-aware behavior distillation approach that combines latent space alignment with noise-invariant auxiliary tasks, enabling effective knowledge transfer from privileged height maps to noisy depth observations. For versatile terrain adaptation, we introduce terrain-specific reward shaping integrated with multi-critic and multi-discriminator learning, where dedicated networks capture the distinct dynamics and motion priors of each terrain type. We validate our approach on two humanoid platforms equipped with different stereo depth cameras. The resulting policy demonstrates robust performance across diverse environments, seamlessly handling extreme challenges such as high platforms and wide gaps, as well as fine-grained tasks including bidirectional long-term staircase traversal.

Dennis Bank, Joost Cordes, Thomas Seel, Simon F. G. Ehlers
总结: 本文提出一种融合深度相机、LiDAR与IMU的CNN-GRU混合自编码器,以生成面向人形机器人行走的稳健中心化高度图,并通过多模态融合和时间上下文显著提升地形重建精度、降低地图漂移。
原摘要

Reliable terrain perception is a critical prerequisite for the deployment of humanoid robots in unstructured, human-centric environments. While traditional systems often rely on manually engineered, single-sensor pipelines, this paper presents a learning-based framework that uses an intermediate, robot-centric heightmap representation. A hybrid Encoder-Decoder Structure (EDS) is introduced, utilizing a Convolutional Neural Network (CNN) for spatial feature extraction fused with a Gated Recurrent Unit (GRU) core for temporal consistency. The architecture integrates multimodal data from an Intel RealSense depth camera, a LIVOX MID-360 LiDAR processed via efficient spherical projection, and an onboard IMU. Quantitative results demonstrate that multimodal fusion improves reconstruction accuracy by 7.2% over depth-only and 9.9% over LiDAR-only configurations. Furthermore, the integration of a 3.2 s temporal context reduces mapping drift.

Ruiqian Nai, Boyuan Zheng, Junming Zhao, ... , Chuan Wen, Yang Gao
总结: HuMI 是一种便携式人形机器人全身操作学习框架,可通过无需机器人的人类示范高效采集全身运动数据,并将其转化为可行的灵巧技能,在多种任务和未知环境中实现比遥操作高 3 倍的数据采集效率和 70% 的成功率。
原摘要

Current approaches for humanoid whole-body manipulation, primarily relying on teleoperation or visual sim-to-real reinforcement learning, are hindered by hardware logistics and complex reward engineering. Consequently, demonstrated autonomous skills remain limited and are typically restricted to controlled environments. In this paper, we present the Humanoid Manipulation Interface (HuMI), a portable and efficient framework for learning diverse whole-body manipulation tasks across various environments. HuMI enables robot-free data collection by capturing rich whole-body motion using portable hardware. This data drives a hierarchical learning pipeline that translates human motions into dexterous and feasible humanoid skills. Extensive experiments across five whole-body tasks--including kneeling, squatting, tossing, walking, and bimanual manipulation--demonstrate that HuMI achieves a 3x increase in data collection efficiency compared to teleoperation and attains a 70% success rate in unseen environments.

Sirui Xu, Samuel Schulter, Morteza Ziyadi, ... , Yu-Xiong Wang, Liangyan Gui
总结: InterPrior 通过大规模模仿预训练、物理扰动数据增强和强化学习微调,学习可泛化的生成式全身控制先验,使类人角色/机器人能在未见目标、初始状态和物体上实现物理一致的人物交互与移动操作。
原摘要

Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors is key to enabling humanoids to compose and generalize loco-manipulation skills across diverse contexts while maintaining physically coherent whole-body coordination. To this end, we introduce InterPrior, a scalable framework that learns a unified generative controller through large-scale imitation pretraining and post-training by reinforcement learning. InterPrior first distills a full-reference imitation expert into a versatile, goal-conditioned variational policy that reconstructs motion from multimodal observations and high-level intent. While the distilled policy reconstructs training behaviors, it does not generalize reliably due to the vast configuration space of large-scale human-object interactions. To address this, we apply data augmentation with physical perturbations, and then perform reinforcement learning finetuning to improve competence on unseen goals and initializations. Together, these steps consolidate the reconstructed latent skills into a valid manifold, yielding a motion prior that generalizes beyond the training data, e.g., it can incorporate new behaviors such as interactions with unseen objects. We further demonstrate its effectiveness for user-interactive control and its potential for real robot deployment.

Harsh Chhajed, Tian Guo
总结: ARBot 是一个开源的高保真机器人遥操作框架,可捕捉并稳定复现自然人类 6-DOF 运动,为增强现实追踪与交互模型提供精确、可重复的地面真值评估,并发布了包含 132 条轨迹的基准数据集。
原摘要

Validating Augmented Reality (AR) tracking and interaction models requires precise, repeatable ground-truth motion. However, human users cannot reliably perform consistent motion due to biomechanical variability. Robotic manipulators are promising to act as human motion proxies if they can mimic human movements. In this work, we design and implement ARBot, a real-time teleoperation platform that can effectively capture natural human motion and accurately replay the movements via robotic manipulators. ARBot includes two capture models: stable wrist motion capture via a custom CV and IMU pipeline, and natural 6-DOF control via a mobile application. We design a proactively-safe QP controller to ensure smooth, jitter-free execution of the robotic manipulator, enabling it to function as a high-fidelity record and replay physical proxy. We open-source ARBot and release a benchmark dataset of 132 human and synthetic trajectories captured using ARBot to support controllable and scalable AR evaluation.

Xiaokang Liu, Zechen Bai, Hai Ci, Kevin Yuchen Ma, Mike Zheng Shou
总结: World-VLA-Loop 通过将状态感知视频世界模型与 VLA 策略进行闭环联合优化,并利用近成功轨迹和失败回放迭代提升模型的动作对齐与策略强化学习效果,从而在极少真实交互下显著提高机器人任务表现。
原摘要

Recent progress in robotic world models has leveraged video diffusion transformers to predict future observations conditioned on historical states and actions. While these models can simulate realistic visual outcomes, they often exhibit poor action-following precision, hindering their utility for downstream robotic learning. In this work, we introduce World-VLA-Loop, a closed-loop framework for the joint refinement of world models and Vision-Language-Action (VLA) policies. We propose a state-aware video world model that functions as a high-fidelity interactive simulator by jointly predicting future observations and reward signals. To enhance reliability, we introduce the SANS dataset, which incorporates near-success trajectories to improve action-outcome alignment within the world model. This framework enables a closed-loop for reinforcement learning (RL) post-training of VLA policies entirely within a virtual environment. Crucially, our approach facilitates a co-evolving cycle: failure rollouts generated by the VLA policy are iteratively fed back to refine the world model precision, which in turn enhances subsequent RL optimization. Evaluations across simulation and real-world tasks demonstrate that our framework significantly boosts VLA performance with minimal physical interaction, establishing a mutually beneficial relationship between world modeling and policy learning for general-purpose robotics. Project page: https://showlab.github.io/World-VLA-Loop/.

Weibin Gu, Chenrui Feng, Lian Liu, ... , Alessandro Rizzo, Guyue Zhou
总结: 本文提出了26克仿蝴蝶机器人 AirPulse,通过柔性低展弦比双翼、低频拍翼和机载闭环分层控制,实现了无尾双翼微型飞行器在强机体振荡下的自主稳定爬升与转向飞行。
原摘要

The flight of biological butterflies represents a unique aerodynamic regime where high-amplitude, low-frequency wingstrokes induce significant body undulations and inertial fluctuations. While existing tailless flapping-wing micro air vehicles typically employ high-frequency kinematics to minimize such perturbations, the lepidopteran flight envelope remains a challenging and underexplored frontier for autonomous robotics. Here, we present \textit{AirPulse}, a 26-gram butterfly-inspired robot that achieves the first onboard, closed-loop controlled flight for a tailless two-winged platform at this scale. It replicates key biomechanical traits of butterfly flight, utilizing low-aspect-ratio, compliant carbon-fiber-reinforced wings and low-frequency flapping that reproduces characteristic biological body undulations. Leveraging a quantitative mapping of control effectiveness, we introduce a hierarchical control architecture featuring state estimator, attitude controller, and central pattern generator with Stroke Timing Asymmetry Rhythm (STAR), which translates attitude control demands into smooth and stable wingstroke timing and angle-offset modulations. Free-flight experiments demonstrate stable climbing and directed turning maneuvers, proving that autonomous locomotion is achievable even within oscillatory dynamical regimes. By bridging biological morphology with a minimalist control architecture, \textit{AirPulse} serves as both a hardware-validated model for decoding butterfly flight dynamics and a prototype for a new class of collision-resilient aerial robots. Its lightweight and compliant structure offers a non-invasive solution for a wide range of applications, such as ecological monitoring and confined-space inspection, where traditional drones may fall short.

Hiroshi Sato, Sho Sakaino, Toshiaki Tsuji
总结: 该论文提出一种无记忆的力生成模仿学习方法,通过从位置轨迹估计力指令并结合稳定的反馈控制,使机器人在未见过的接触丰富轨迹上也能生成有效力控制命令。
原摘要

In contact-rich tasks, while position trajectories are often easy to obtain, appropriate force commands are typically unknown. Although it is conceivable to generate force commands using a pretrained foundation model such as Vision-Language-Action (VLA) models, force control is highly dependent on the specific hardware of the robot, which makes the application of such models challenging. To bridge this gap, we propose a force generative model that estimates force commands from given position trajectories. However, when dealing with unseen position trajectories, the model struggles to generate accurate force commands. To address this, we introduce a feedback control mechanism. Our experiments reveal that feedback control does not converge when the force generative model has memory. We therefore adopt a model without memory, enabling stable feedback control. This approach allows the system to generate force commands effectively, even for unseen position trajectories, improving generalization for real-world robot writing tasks.