Papers for 2026-09-25

10 papers
Al Jaber Mahmud, Shuai Li, Xuan Wang
总结: 联合选择支撑接触与全身配置以平衡任务能力。
方法: CTCS筛选分组候选并用敏感性预测能力。
证据: 局部敏感性分析减少重复全身优化成本。
为什么适合我: 直接服务腿足loco-manipulation接触决策。
推荐理由: 直接命中腿足 loco-manipulation:在接触丰富任务中联合选择环境支撑接触与全身构型,并权衡残差扳手、末端可达与基座机动。
原摘要

In this paper, we study the joint selection of an environmental support contact and a whole-body configuration for a prescribed loco-manipulation task. A contact may provide greater physical support while restricting the motion required for the task. We formulate this problem through three capability measures: residual wrench, end-effector reach, and base mobility available after satisfying the task requirements, and we balance them against contact acquisition cost. Evaluating these capabilities for every candidate requires repeated whole-body optimizations. To reduce this computational cost, we propose Capability-Tradeoff Contact Selection (CTCS). CTCS screens candidates for contact and task feasibility, groups similar candidates within each surface, and predicts their capabilities from exact anchor evaluations using local sensitivity analysis. It checks these predictions through selective exact evaluations, ranks candidates by capability, and evaluates a shortlist exactly for final selection. We evaluate CTCS in simulations and hardware experiments using a Unitree Go2 quadruped with an AgileX NERO arm across $392$ task conditions with nine available support surfaces. Results show that CTCS outperforms ground-only and fixed-contact support, as it can select support surfaces that provide favorable capability trade-offs for the task. Compared with evaluating every candidate exactly, CTCS achieves approximately $3\times$ speedup while closely matching the resulting mean objective value.

DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models

5.0/5 很相关 裁判分 9.0 Top pick 当日相对
Yohan Choi, Min-Jun Kim, Jin-Sung Kim, Yong-Jae Kim, Youn-Hee Han
总结: 深度去噪世界模型实现噪声鲁棒四足跑酷。
方法: 噪声深度编码干净重建并对比对齐潜状态。
证据: 鲁棒性嵌入学习管道消除部署滤波依赖。
为什么适合我: 契合复杂地形感知运动与跑酷智能。
推荐理由: 直接命中腿足感知运动与跑酷:四足跑酷把深度噪声鲁棒性做进世界模型,对应深度/高程感知与 sim-to-real 部署。
原摘要

Vision-based legged locomotion methods assume clean depth at training time and rely on hand-tuned post-processing filters at deployment. However, filter parameters are rarely disclosed, hindering reproducibility, and performance degrades substantially when depth noise is left unaddressed. Building noise robustness directly into the learning pipeline would eliminate this dependency. While such robustness has been explored for proprioceptive inputs, analogous approaches for depth perception remain largely absent in legged locomotion. We propose DAWN (Denoising and Alignment in World models for Noise-robustness), a noise-robust perception framework for legged locomotion, which builds noise robustness directly into a world model via two modifications: (1) feeding noisy depth to the encoder while keeping clean depth as the reconstruction target, forcing the model to implicitly denoise its input; and (2) applying contrastive learning to align the latent states of noisy and clean depth. Importantly, DAWN is not tied to a specific noise model, requiring no manual tuning to the noise distribution at deployment. Furthermore, it incurs no additional inference cost over existing world model-based methods. Without any manual filter calibration -- relying solely on the learned noise-robust representation -- DAWN achieves zero-shot quadruped parkour on a Unitree Go1: traversing stairs up to 18 cm, clearing gaps up to 70 cm, and mounting steps up to 45 cm from raw depth observations. Ablation studies show that denoising and contrastive alignment contribute at complementary levels -- reconstruction and representation, respectively -- and yield additive gains when combined. Videos and code are available at: https://dawn-parkour.github.io/

Tara Sadjadpour, Siming He, C. K. Wolfe, ... , Claire Tomlin, Jitendra Malik
总结: 三阶段框架将手物交互转为零样本视觉策略。
方法: 形态优化重定向加残差RL再蒸馏成策略。
证据: 接触F1超强基线8至28点并提升成功率。
为什么适合我: 接触感知重定向启发人体动作全身迁移。
推荐理由: 最接近人体动作重定向与模仿:接触保持的手部重定向、残差 RL 与 sim-to-real 视觉运动策略,但是手而非全身。
原摘要

Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO outperforms five baselines in contact F1, improving on the strongest ones by 8 to 28 points, while improving the success rate of downstream dynamic retargeting by as much as 35 points. On a Sharpa hand, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects spanning 10 categories. Project page: https://morphometricimitation.github.io

Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching

3.0/5 一般 裁判分 5.0 Candidate 当日相对
Guorui Pei, Jinsong Wu, Songyuan Su, ... , Bin Liu, Peng Zhou
总结: 结果敏感运动搜索实现冲击感知灵巧抓取。
方法: 学习成功窗口流形并测地搜索精炼示范。
证据: 干预结果敏感性指导针对性示范构建。
为什么适合我: 冲击接触场景下感知运动控制高度相关。
推荐理由: 接触敏感的灵巧抓取,RL 教师加模仿学生与接触过渡相关,但不是全身运动跟踪或腿足 loco-manipulation。
原摘要

Skilled humans can catch fast-moving objects softly by coordinating interception, velocity matching, and follow-through to mitigate impact. Learning such impact-aware catching with reinforcement learning (RL), however, is challenging, as the policy must achieve reliable interception and grasping while regulating the sensitive transition into contact. Moreover, even a capable privileged-state RL teacher may not provide ideal demonstrations for a deployable imitation-learning (IL) student: teacher failures limit task coverage, while small variations in pre-contact motion can produce substantially different impact and grasping outcomes. We characterize this phenomenon through interventional outcome sensitivity and introduce the outcome-sensitive window (OSW) to guide targeted demonstration construction. Building on this formulation, we propose Outcome-Sensitive Motion Search, which learns a task-conditioned manifold of successful OSW motions and performs local geodesic search to refine successful teacher rollouts and repair task conditions where the teacher fails. We then validate candidate motions through complete rollouts under a calibrated IL-student action-error model and retain only successful executions as demonstrations. Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.

Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
总结: 人类引导残差强化学习实现高效灵巧操作。
方法: 冻结模仿策略上学习残差纠正并塑形奖励。
证据: 仅20初始示范在五接触任务高效学习。
为什么适合我: 残差RL可精炼模仿用于全身控制部署。
推荐理由: 接触丰富灵巧操作上的人类残差强化学习,方法靠近模仿加 RL,但对象不是人形全身或腿足场景交互。
原摘要

Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.

EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies

3.0/5 一般 裁判分 5.0 Candidate 当日相对
Hanbit Oh, Yukiyasu Domae, Takuma Yagi
总结: 将人类操作节奏迁移到机器人模仿策略。
方法: 对齐相位估计相对节奏并重定时机器人示范。
证据: 人类示范提供任务适宜相位节奏监督。
为什么适合我: 人体动作节奏迁移支持可跟踪全身控制。
推荐理由: 用人体示范给操作策略提供阶段节奏,只擦边人体到机器人的模仿,不是全身重定向、运动跟踪或 loco-manipulation。
原摘要

Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.

Jun Hu, Sihan Chen, Kosta Jovanovic, ... , Jia Pan, Peng Zhou
总结: 电机电流对齐实现单臂链接接触丰富举升。
方法: 分阶段特权策略生成示范并因果映射电流。
证据: 统一学生策略从对齐示范学习长时程任务。
为什么适合我: 全身臂接触操作启发腿足loco-manipulation。
推荐理由: 单臂多连杆接触举起大物体,接近接触丰富场景交互与 sim-to-real,但不是腿足全身 loco-manipulation。
原摘要

Most robots manipulate objects solely with their end effectors, whereas humans flexibly leverage different body parts, such as the forearm and elbow, especially when handling oversized objects. Learning such whole-arm manipulation is chal-lenging due to long-horizon sparse rewards, limited contact sens-ing, and the sim-to-real gap in contact and actuator dynamics. To address these challenges, we propose Current-Aligned Link Manipulation, a framework for learning long-horizon contact-rich manipulation using motor current as joint load related feedback. Three stage-specific policies first learn repositioning, grasping, and lifting using privileged simulation information, and a stage router sequences them to generate complete task demonstrations. For sim-to-real transfer, a causal current mapper predicts physical motor current from simulated joint histories, aligning the actuator current observation between simulation and hardware. A unified student policy then learns from these demonstrations using only deployable sensor observations and is further refined with DAgger. The task policies are trained entirely in simulation, and the final student is deployed on hardware. Experiments demonstrate 76.2% (762/1000 trials) complete-task success in simulation and 73.3% success (22/30 trials) on the physical robot for sequential oversized-object lifting.

Online Sim-to-Real Adaptation via Closed-Loop System Modeling

2.0/5 偏低 裁判分 8.0 Strong 当日相对
Yuhao Huang, Samuel A. Moore, Boyuan Chen
总结: 闭环系统建模实现在线仿真到现实适应。
方法: 元训练闭环模型微调后适应参考命令。
证据: 有限真实交互快速微调提升跟踪精度。
为什么适合我: 仿真到真机适应支持全身控制上机部署。
推荐理由: 在线仿真到真实的闭环指令自适应与迁移主题相邻,但未体现腿式或人形全身控制器。
原摘要

Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Yangang Ren, Yujie Yan, Zirui Li, ... , Xuesong Tian, Chen Lv
总结: 从混合质量部署经验学习机器人操作策略。
方法: 预测动作块批评家增强长时程价值估计。
证据: 批评家将混合经验转为有效模仿学习信号。
为什么适合我: 部署经验学习可精炼真机全身运动策略。
推荐理由: 通用操作的部署后模仿/离线强化学习,不涉及人形或腿足全身控制、跑酷与 loco-manipulation。
原摘要

Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.

Tomohiro Motoda, Masaki Murooka, Keisuke Shirai, ... , Hugo Duarte, Yukiyasu Domae
总结: 自监督锚定指尖传感到本体感觉用于模仿。
方法: PROPRA预训练将传感锚定到本体与主动动作。
证据: 解决稀疏相位依赖信号难从示范利用问题。
为什么适合我: 接触传感锚定启发全身接触丰富感知运动。
推荐理由: 指尖触觉/接近觉的模仿学习,属于局部操作感知,不接近全身控制、动作重定向或腿足运动。
原摘要

Robotic imitation learning often relies on external cameras, yet local interaction cues such as object proximity, contact onset, and grasp state are difficult to observe near the fingertips because of occlusion and limited temporal resolution. We study how to effectively incorporate complementary fingertip sensing into imitation learning using pressure-sensitive tactile and reflective proximity sensors, along with pretrained sensor encoders. The two modalities provide information at different manipulation phases: proximity sensing is informative before contact, whereas tactile sensing becomes informative after contact. However, naively adding these signals to a policy does not consistently improve performance and can even underperform vision-only policies, suggesting that sparse, phase-dependent sensor signals are difficult to exploit from limited demonstrations. We therefore propose a proprioception-anchored pretraining method, PROprioceptive-and-PRoactive Anchoring (PROPRA), which independently aligns each fingertip sensor history with proprioceptive and action segments. This provides a continuously available sensorimotor reference, allowing each sensor to be aligned independently during its informative phases. Experiments on real-world manipulation tasks show that our pretraining method improves average success rates over vision-only policies and image-anchored pretraining baselines. Representation analysis further shows that it preserves richer information about pre-contact states, enabling more effective use of complementary fingertip sensing. Please refer to our project page: https://tomohiromotoda.github.io/nia.propra/