Papers for 2026-03-08

10 papers
Dayang Liang, Yuhang Lin, Xinzhe Liu, ... , Yunlong Liu, Chenjia Bai
总结: InterReal 是一个统一的基于物理的模仿学习框架,通过带手-物接触约束的 HOI 数据增强和自动奖励学习,使类人机器人能更稳定、准确地学习并在真实环境中执行抓箱、推箱等人-物交互技能。
原摘要

Interaction is one of the core abilities of humanoid robots. However, most existing frameworks focus on non-interactive whole-body control, which limits their practical applicability. In this work, we develop InterReal, a unified physics-based imitation learning framework for Real-world human-object Interaction (HOI) control. InterReal enables humanoid robots to track HOI reference motions, facilitating the learning of fine-grained interactive skills and their deployment in real-world settings. Within this framework, we first introduce a HOI motion data augmentation scheme with hand-object contact constraints, and utilize the augmented motions to improve policy stability under object perturbations. Second, we propose an automatic reward learner to address the challenge of large-scale reward shaping. A meta-policy guided by critical tracking error metrics explores and allocates reward signals to the low-level reinforcement learning objective, which enables more effective learning of interactive policies. Experiments on HOI tasks of box-picking and box-pushing demonstrate that InterReal achieves the best tracking accuracy and the highest task success rate compared to recent baselines. Furthermore, we validate the framework on the real-world robot Unitree G1, which demonstrates its practical effectiveness and robustness beyond simulation.

Yufei Liu, Xieyuanli Chen, Hainan Pan, ... , Zhiwen Zeng, Huimin Lu
总结: GeoLoco利用冻结的具备尺度感知能力的视觉基础模型从单目RGB中提取3D几何先验,并通过本体感知查询的跨注意力与辅助几何对齐训练,实现了仅在仿真中训练即可零样本迁移到Unitree G1人形机器人并稳健通过复杂地形的RGB-only运动控制。
原摘要

The prevailing paradigm of perceptive humanoid locomotion relies heavily on active depth sensors. However, this depth-centric approach fundamentally discards the rich semantic and dense appearance cues of the visual world, severing low-level control from the high-level reasoning essential for general embodied intelligence. While monocular RGB offers a ubiquitous, information-dense alternative, end-to-end reinforcement learning from raw 2D pixels suffers from extreme sample inefficiency and catastrophic sim-to-real collapse due to the inherent loss of geometric scale. To break this deadlock, we propose GeoLoco, a purely RGB-driven locomotion framework that conceptualizes monocular images as high-dimensional 3D latent representations by harnessing the powerful geometric priors of a frozen, scale-aware Visual Foundation Model (VFM). Rather than naive feature concatenation, we design a proprioceptive-query multi-head cross-attention mechanism that dynamically attends to task-critical topological features conditioned on the robot's real-time gait phase. Crucially, to prevent the policy from overfitting to superficial textures, we introduce a dual-head auxiliary learning scheme. This explicit regularization forces the high-dimensional latent space to strictly align with the physical terrain geometry, ensuring robust zero-shot sim-to-real transfer. Trained exclusively in simulation, GeoLoco achieves robust zero-shot transfer to the Unitree G1 humanoid and successfully negotiates challenging terrains.

Zhaoyang Xiang, Upama Pant, Ayonga Hereid
总结: 本文提出一种基于机载感知的混合整数模型预测控制方法,可在离散可踏区域上联合规划人形机器人落脚点与步时,并通过DCM可捕获性约束和步内重规划实现安全、鲁棒且毫秒级求解的动态行走。
原摘要

Many real-world walking scenarios contain obstacles and unsafe ground patches (e.g., slippery or cluttered areas), leaving a disconnected set of admissible footholds that can be modeled as stepping-stone-like regions. We propose an onboard, perceptive mixed-integer model predictive control framework that jointly plans foot placement and step duration using step-to-step Divergent Component of Motion (DCM) dynamics. Ego-centric depth images are fused into a probabilistic local heightmap, from which we extract a union of convex steppable regions. Region membership is enforced with binary variables in a mixed-integer quadratic program (MIQP). To keep the optimization tractable while certifying safety, we embed capturability bounds in the DCM space: a lateral one-step condition (preventing leg crossing) and a sagittal infinite-step bound that limits unstable growth. We further re-plan within the step by back-propagating the measured instantaneous DCM to update the initial DCM, improving robustness to model mismatch and external disturbances. We evaluate the approach in simulation on Digit on randomized stepping-stone fields, including external pushes. The planner generates terrain-aware, dynamically consistent footstep sequences with adaptive timing and millisecond-level solve times.

Hangjun Liu, Jiarui Geng, Jinxuan Ding, ... , Grigoriy Blekherman, Baxi Chong
总结: 该论文通过比较蜈蚣生物力学与可调腿长机器人实验,提出统一侧弯与滚转的波动式自翻正框架,揭示腿长和腿数等形态参数决定自翻正控制策略,并指出过长肢体会使细长多足机器人难以可靠翻身。
原摘要

Centipede-like robots offer unique locomotion advantages due to their small cross-sectional area for accessing confined spaces, and their redundant legs enhance robustness in cluttered environments such as search-and-rescue and pipe inspection. However, elongated robots are particularly vulnerable to tipping over when climbing large obstacles, making reliable self-righting essential for field deployment. Self-righting strategies for elongate, multi-legged systems remain poorly understood. In this study, we conduct a comparative biomechanics and robophysical investigation to address three key questions: (1) What self-righting strategies are effective for elongate, many-legged systems? (2) How should these strategies depend on morphological parameters such as leg length and leg number? (3) Is there a morphological limit beyond which reliable self-righting becomes infeasible? We compare two biological exemplars: Scolopendra subspinipes (short legs) and Scutigera coleoptrata (house centipedes with long legs). Scolopendra subspinipes reliably self-rights both during aerial phases and through ground-assisted self-righting, whereas house centipedes rely predominantly on aerial reorientation and struggle to generate effective self-righting torques during ground contact. Motivated by these observations, we construct a parameterized space of bio-inspired self-righting strategies and develop an elongate robot with adjustable leg lengths. Systematic experiments reveal that increasing leg length necessitates a shift in control strategy to prevent torque over-concentration in mid-body actuators, and we identify a critical limb-length threshold above which robust self-righting becomes challenging. These results establish morphology-strategy coupling principles for self-righting in elongate robots and provide design guidelines for centipede-like systems operating in uncertain terrain.

Zihang You, Xianlian Zhou
总结: 本文提出并验证了一种基于强化学习的仿真训练外骨骼控制框架,可通过髋膝运动学预测辅助力矩以降低人体关节力矩,并在公开步态数据上显示出尤其在髋关节力矩时序上的较强仿真到真实一致性,但在高速、陡坡和关节功率匹配方面仍存在挑战。
原摘要

Data-driven joint-moment predictors offer a scalable alternative to laboratory-based inverse-dynamics pipelines for biomechanics estimation and exoskeleton control. Meanwhile, physics-based reinforcement learning (RL) enables simulation-trained controllers to learn dynamics-aware assistance strategies without extensive human experimentation. However, quantitative verification of simulation-trained exoskeleton torque predictors, and their impact on human joint power injection, remains limited. This paper presents (1) an RL framework to learn exoskeleton assistance policies that reduce biological joint moments, and (2) a validation pipeline that verifies the trained control networks using an open-source gait dataset through inference and comparison with biological joint moments. Simulation-trained multilayer perceptron (MLP) controllers are developed for level-ground and ramp walking, mapping short-horizon histories of bilateral hip and knee kinematics to normalized assistance torques. Results show that predicted assistance preserves task-intensity trends across speeds and inclines. Agreement is particularly strong at the hip, with cross-correlation coefficients reaching 0.94 at 1.8 m/s and 0.98 during 5° decline walking, demonstrating near-matched temporal structure. Discrepancies increase at higher speeds and steeper inclines, especially at the knee, and are more pronounced in joint power comparisons. Delay tuning biases assistance toward greater positive power injection; modest timing shifts increase positive power and improve agreement in specific gait intervals. Together, these results establish a quantitative validation framework for simulation-trained exoskeleton controllers, demonstrate strong sim-to-data consistency at the torque level, and highlight both the promise and the remaining challenges for sim-to-real transfer.

Chen Xiong, Cheng Wang, Yuhang Liu, Zirui Wu, Ye Tian
总结: 该论文提出一种行为引导的紧急变道高风险场景生成方法,通过优化序列生成对抗网络学习少量真实应急变道行为,并结合递归PPO与模型预测控制,高效生成物理真实的碰撞风险仿真场景。
原摘要

In contemporary autonomous driving testing, virtual simulation has become an important approach due to its efficiency and cost effectiveness. However, existing methods usually rely on reinforcement learning to generate risky scenarios, making it difficult to efficiently learn realistic emergency behaviors. To address this issue, we propose a behavior guided method for generating high risk lane change scenarios. First, a behavior learning module based on an optimized sequence generative adversarial network is developed to learn emergency lane change behaviors from an extracted dataset. This design alleviates the limitations of existing datasets and improves learning from relatively few samples. Then, the opposing vehicle is modeled as an agent, and the road environment together with surrounding vehicles is incorporated into the operating environment. Based on the Recursive Proximal Policy Optimization strategy, the generated trajectories are used to guide the vehicle toward dangerous behaviors for more effective risk scenario exploration. Finally, the reference trajectory is combined with model predictive control as physical constraints to continuously optimize the strategy and ensure physical authenticity. Experimental results show that the proposed method can effectively learn high risk trajectory behaviors from limited data and generate high risk collision scenarios with better efficiency than traditional methods such as grid search and manual design.

Nico Messikommer, Jiaxu Xing, Leonard Bauersfeld, ... , Elie Aljalbout, Davide Scaramuzza
总结: 本文提出一种近似模仿学习框架,使四旋翼仅依靠单个事件相机和端到端神经网络即可在杂乱环境中高速飞行,并通过避免训练时渲染昂贵事件数据实现高效仿真训练,在仿真和真实测试中均优于基线且实飞速度达9.8 m/s。
原摘要

Event cameras offer high temporal resolution and low latency, making them ideal sensors for high-speed robotic applications where conventional cameras suffer from image degradations such as motion blur. In addition, their low power consumption can enhance endurance, which is critical for resource-constrained platforms. Motivated by these properties, we present a novel approach that enables a quadrotor to fly through cluttered environments at high speed by perceiving the environment with a single event camera. Our proposed method employs an end-to-end neural network trained to map event data directly to control commands, eliminating the reliance on standard cameras. To enable efficient training in simulation, where rendering synthetic event data is computationally expensive, we propose Approximate Imitation Learning, a novel imitation learning framework. Our approach leverages a large-scale offline dataset to learn a task-specific representation space. Subsequently, the policy is trained through online interactions that rely solely on lightweight, simulated state information, eliminating the need to render events during training. This enables the efficient training of event-based control policies for fast quadrotor flight, highlighting the potential of our framework for other modalities where data simulation is costly or impractical. Our approach outperforms standard imitation learning baselines in simulation and demonstrates robust performance in real-world flight tests, achieving speeds up to 9.8 ms-1 in cluttered environments.

Junkun Jiang, Jie Chen, Ho Yin Au, Jingyu Xiang
总结: 该论文提出基于扩散模型的Masked Motion Diffusion Model,通过高效的运动学注意力聚合学习上下文自适应运动先验,从不完整或低置信度数据中实现更准确的3D动作补全、修复与插帧。
原摘要

Vision-based motion capture solutions often struggle with occlusions, which result in the loss of critical joint information and hinder accurate 3D motion reconstruction. Other wearable alternatives also suffer from noisy or unstable data, often requiring extensive manual cleaning and correction to achieve reliable results. To address these challenges, we introduce the Masked Motion Diffusion Model (MMDM), a diffusion-based generative reconstruction framework that enhances incomplete or low-confidence motion data using partially available high-quality reconstructions within a Masked Autoencoder architecture. Central to our design is the Kinematic Attention Aggregation (KAA) mechanism, which enables efficient, deep, and iterative encoding of both joint-level and pose-level features, capturing structural and temporal motion patterns essential for task-specific reconstruction. We focus on learning context-adaptive motion priors, specialized structural and temporal features extracted by the same reusable architecture, where each learned prior emphasizes different aspects of motion dynamics and is specifically efficient for its corresponding task. This enables the architecture to adaptively specialize without altering its structure. Such versatility allows MMDM to efficiently learn motion priors tailored to scenarios such as motion refinement, completion, and in-betweening. Extensive evaluations on public benchmarks demonstrate that MMDM achieves strong performance across diverse masking strategies and task settings. The source code is available at https://github.com/jjkislele/MMDM.

Jingzehua Xu, Guanwen Xie, Jiwei Tang, Shuai Zhang, Xiaofan Li
总结: 本文综述从“约束耦合”视角指出,水下自主机器人在真实海洋中的可靠智能必须将水动力不确定性、部分可观测、通信受限与能量稀缺等因素纳入感知—规划—控制闭环统一建模,并提出面向物理世界模型、可认证学习控制、通信感知协同和部署导向设计的未来方向。
原摘要

Autonomous underwater robots are increasingly deployed for environmental monitoring, infrastructure inspection, subsea resource exploration, and long-horizon exploration. Yet, despite rapid advances in learning-based planning and control, reliable autonomy in real ocean environments remains fundamentally constrained by tightly coupled physical limits. Hydrodynamic uncertainty, partial observability, bandwidth-limited communication, and energy scarcity are not independent challenges; they interact within the closed perception-planning-control loop and often amplify one another over time. This Review develops a constraint-coupled perspective on underwater embodied intelligence, arguing that planning and control must be understood within tightly coupled sensing, communication, coordination, and resource constraints in real ocean environments. We synthesize recent progress in reinforcement learning, belief-aware planning, hybrid control, multi-robot coordination, and foundation-model integration through this embodied perspective. Across representative application domains, we show how environmental monitoring, inspection, exploration, and cooperative missions expose distinct stress profiles of cross-layer coupling. To unify these observations, we introduce a cross-layer failure taxonomy spanning epistemic, dynamic, and coordination breakdowns, and analyze how errors cascade across autonomy layers under uncertainty. Building on this structure, we outline research directions toward physics-grounded world models, certifiable learning-enabled control, communication-aware coordination, and deployment-aware system design. By internalizing constraint coupling rather than treating it as an external disturbance, underwater embodied intelligence may evolve from performance-driven adaptation toward resilient, scalable, and verifiable autonomy under real ocean conditions.

Tongqing Chen, Hang Wu, Jiasen Wang, ... , Zhu Jin, Lu Fang
总结: SuperSuit通过统一的双模态同构运动接口将人类步行映射为移动底盘控制、将可穿戴手臂映射为机械臂关节轨迹,从而高效采集可直接混合训练的长时程移动操作示范数据,并在真实任务中显著提升数据采集吞吐量和策略性能。
原摘要

High-quality, long-horizon demonstrations are essential for embodied AI, yet acquiring such data for tightly coupled wheeled mobile manipulators remains a fundamental bottleneck. Unlike fixed-base systems, mobile manipulators require continuous coordination between $SE(2)$ locomotion and precise manipulation, exposing limitations in existing teleoperation and wearable interfaces. We present \textbf{SuperSuit}, a bimodal data acquisition framework that supports both robot-in-the-loop teleoperation and active demonstration under a shared kinematic interface. Both modalities produce structurally identical joint-space trajectories, enabling direct data mixing without modifying downstream policies. For locomotion, SuperSuit maps natural human stepping to continuous planar base velocities, eliminating discrete command switches. For manipulation, it employs a strictly isomorphic wearable arm in both modes, while policy training is formulated in a shift-invariant delta-joint representation to mitigate calibration offsets and structural compliance without inverse kinematics. Real-world experiments on long-horizon mobile manipulation tasks show 2.6$\times$ higher demonstration throughput in active mode compared to a teleoperation baseline, comparable policy performance when substituting teleoperation data with active demonstrations at fixed dataset size, and monotonic performance improvement as active data volume increases. These results indicate that consistent kinematic representations across collection modalities enable scalable data acquisition for long-horizon mobile manipulation.