Papers for 2026-09-28

10 papers
Shuliang He, Ruiyan Xu, Bo Yue, ... , Wei-Shi Zheng, Guiliang Liu
总结: 从第一人称视频蒸馏交互先验实现可泛化全身操作。
方法: 视觉语言导航、闭环姿态校准与在线感知协调三阶段。
证据: 单次演示指定技能,在线视觉触觉反馈适应新物体姿态。
为什么适合我: 直接支撑人形loco-manipulation与接触丰富全身控制。
推荐理由: 移动人形全身操作:自我中心视频交互先验、上下身协同与闭环姿态校准,直接对应全身 loco-manipulation 与人体交互重定向。
原摘要

Mobile humanoid manipulation requires both reaching a usable workspace and preserving precise hand-object interactions as object poses and contact conditions change. Learning these behaviors from limited task-specific data remains challenging. To bridge this gap, we introduce Praxis, a whole-body manipulation framework that combines physical interaction priors from one-shot egocentric video demonstrations with closed-loop posture calibration and online perception. The framework coordinates three stages: vision-language-guided navigation toward target objects, closed-loop posture calibration to align the arm-hand workspace, and dexterous manipulation with synchronized upper- and lower-body control. Online visual feedback re-grounds demonstrated interaction geometry under new object poses and scene configurations, while tactile feedback adapts hand motions to actual contact conditions. Each manipulation skill is specified by one human demonstration, without task-specific manipulation-policy retraining. Experiments across five long-horizon manipulation tasks demonstrate spatial, visual, and cross-object generalization, as well as recovery from external physical disturbances across all three stages.

Dyuman Aditya, Jin Cheng, Clemens Schwarke, ... , Stelian Coros, Gabriele Fadini
总结: 捆绑接触梯度稳定可微仿真用于动态任务部署。
方法: 接触局部随机平滑框架处理刚性接触高方差梯度。
证据: 提升动态人形运动策略向真实世界的迁移保真度。
为什么适合我: 支持人形接触丰富动态运动的可微优化与真机迁移。
推荐理由: 针对可部署动态人形的接触可微仿真与策略梯度稳定,紧贴全身接触控制与 sim-to-real。
原摘要

Differentiable simulation provides analytic gradients of robot dynamics, enabling fast and sample-efficient first-order policy optimization. However, obtaining smooth and informative gradients through rigid-body contact typically requires softened contact models, often at the expense of physical fidelity and thereby limiting learned policies largely to simulation. This trade-off becomes particularly consequential for dynamic humanoid motions, where accurate contact dynamics are critical for transferring policies to the real world. Increasing contact stiffness in rigid-body simulation improves the fidelity of interactions, but also makes the dynamics increasingly sensitive to small state perturbations, producing high-variance gradients that can destabilize first-order policy learning. To address this, we propose \emph{Bundled Contact Gradients (BCG)}, a contact-local randomized smoothing framework for differentiable policy learning. When stiff contact is detected, our method evaluates a local bundle of randomized perturbation rollouts around the stiff contact configuration and aggregates their gradient signal thereby reducing gradient variance. We demonstrate the effectiveness of our method by successfully training and transferring dynamic motions zero-shot onto a real-world Unitree G1 humanoid platform. Videos and supplementary information can be found at https://bundledcontactgradients.github.io/

Giammarco Caroleo, Timothée Mahamoodally, Matteo Manzardo, ... , Perla Maiolino, Maurice Fallon
总结: 分布式低成本ToF实现四足近场地形映射与避障。
方法: 为ANYmal设计分布式ToF架构并与深度相机对比。
证据: 低分辨率高噪声下仍支持厘米级可靠感知运动。
为什么适合我: 助力腿足机器人复杂地形低成本感知与脚步规划。
推荐理由: 四足近场地形感知、落足与避障,直接对应腿足感知运动;ToF 替代深度/高程图,而非 SLAM 或激光里程计。
原摘要

Quadruped robots typically rely on depth cameras and LiDAR sensors to map their local environment. However, these sensors have limited close-range coverage, are relatively expensive, and consume significant power. This study investigates whether distributed Time-of-Flight (ToF) sensors can serve as a low-cost alternative to depth cameras for near-field terrain mapping for locomotion and local navigation. We designed a distributed ToF sensing architecture for the ANYbotics ANYmal quadruped, assessed its environment reconstruction accuracy, and benchmarked it against depth cameras for terrain mapping and obstacle avoidance. Distributing these sensors around the robot can also avoid the blind spots of traditional sensors. Our results show that, despite their low resolution and higher measurement noise, distributed ToF sensors can support reliable perceptual locomotion with centimeter-level local mapping accuracy. The proposed sensing strategy provides sufficient geometric information for near-field obstacle avoidance and footstep planning, at substantially lower cost, energy consumption, and system complexity than depth cameras.

Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
总结: 运动学感知MMDiT框架生成协调人体运动。
方法: 共享多模态注意力加流匹配与旋转运动学监督。
证据: 两阶段课程解决直接MMDiT导致的抖动不协调。
为什么适合我: 利于人体动作生成可重定向可跟踪的全身扩散控制。
推荐理由: 运动学感知的扩散式人体动作生成,接近物理角色运动先验,但未涉及跟踪、重定向或真机全身控制。
原摘要

Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing multimodal diffusion transformers (MMDiT), which have shown effective joint text--visual modeling in vision generation, into HMG. However, we find that articulated motion is temporally coherent but weakly correlated across joints, in which directly applying an MMDiT with flow matching produces poorly coordinated and jerky motion. In this work, we propose Timo, a novel kinematics-aware MMDiT framework tailored for HMG. Timo combines fully shared multimodal attention for bidirectional text--motion modeling with flow matching, geometric and rotational-kinematics supervision that compares actual rotations and their changes over time, and a two-stage curriculum progressing from broad motion learning to detailed caption alignment. Further, we construct a benchmark of $40{,}025$ held-out clips from six public datasets spanning diverse actions, assessing six complementary dimensions under a common evaluator and scoring protocol. Our model substantially outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Remarkably, Timo surpasses Kimodo on five of six dimensions, achieving a $40.8$% relative improvement in the average benchmark score. Project page: https://kyfafyd.wang/projects/timo. Demo page: https://timo.kyfafyd.wang.

Alexander Alexiev, Tzu-Yuan Lin, Sang Min Kim, ... , Yonghyeon Lee, Sangbae Kim
总结: 仅用本体感觉学习拟人手盲抓取反射。
方法: 分离臂运动与手部RL策略,仅靠本体反馈抓稳。
证据: 仿真硬件验证多样物体稳健抓取并可组合臂控。
为什么适合我: 支撑接触丰富场景感知运动分离与loco-manipulation。
推荐理由: 仿人灵巧手的本体感觉盲抓反射与接触 RL,接近手部接触技能,但缺少全身协调与腿足运动。
原摘要

In this work we study if a robotic hand using proprioception alone can grasp diverse objects with no visual observation. We present a modular dexterous grasping architecture that separates global arm motion from local contact control. An independently controlled arm guides the hand toward the object, while a reinforcement learning policy grasps and stabilizes it using only hand proprioceptive feedback. We call this \textit{a blind grasp reflex}: grasping without images, object poses, or geometric observations. A learned stable-grasp score determines when the object is securely held, allowing the arm to begin post-grasp manipulation. This separation makes grasping a reusable hand-level skill that can be combined with independently designed arm controllers for various manipulation tasks. Experiments in simulation and on hardware demonstrate robust blind grasping across diverse objects and seamless composition with a range of arm controllers. Moreover, despite never observing contact geometry, the learned grasp score closely aligns with an independent physics-based measure of grasp stability. The resulting approach follows a simple principle: see to reach, feel to grasp. Project page: https://blindgraspreflex.github.io.

Namiko Saito, Kinam Kim, Heecheol Kim, Katsushi Ikeuchi, Yasuyuki Matsushita
总结: 仿真潜在条件残差RL增强冻结VLA精确执行。
方法: 用VLA潜在表示条件残差并学习轻量sim-real映射。
证据: 四个接触丰富任务与两个VLA骨干验证有效迁移。
为什么适合我: 结合RL与VLA提升接触操作真机部署与全身控制。
推荐理由: 冻结 VLA 的仿真残差 RL 与 sim-to-real,属于通用接触操作,不是人形/腿足全身运动。
原摘要

Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.

Hayato Takahashi, Ryoga Oishi, Yuki Kasuga, Toshiaki Tsuji
总结: 克隆刚度与平衡点参数实现接触丰富阻抗控制。
方法: 从双边遥操作经粒子滤波提取参数无需力传感器。
证据: 擦拭任务维持一致接触力优于轨迹模仿基线。
为什么适合我: 模仿生物力学先验利于接触场景全身控制泛化。
推荐理由: 固定机械臂阻抗克隆与擦拭接触,仅有接触丰富操作,无全身或腿足运动。
原摘要

Contact-rich manipulation requires robots to regulate force against surfaces whose geometry deviates unpredictably from training conditions. Trajectory-based imitation learning, which reproduces observable outputs, breaks down under such shifts. We propose Impedance Cloning, which instead imitates the biomechanical priors that generate motion -- the stiffness and equilibrium point -- and thereby passively absorbs contact uncertainty. Because these parameters encode intent rather than outcome, they generalize across surface geometries where trajectory reproduction does not. We extract them from bilateral teleoperation demonstrations via a particle filter without force/torque sensors and evaluate the framework on two CRANE-X7 manipulators. In a wiping task with joint-space actions, the trajectory-based baseline loses contact below -6 cm, whereas the proposed method maintains a consistent 4-5 N contact force above -6 cm, with a gradual decrease below; with Cartesian-space actions, its force-height slope over 0 to +8 cm is 0.13 +/- 0.03 N/cm, versus 0.34-0.83 N/cm for fixed-impedance baselines. In a pick-and-place task with 10 diverse cups (100 trials), the proposed method succeeds in 84 trials, outperforming the fixed-impedance baseline (74/100) and performing comparably to a variable impedance control baseline (82/100) with one demonstration instead of ten. In a grasping task, the representation reduces torque tracking error with both ILBiT and Mamba backbones, confirming its generality across architectures.

Chen-Chieh Liao, Yichen Peng, Yiyi Cai, ... , Hideki Koike, Shuichi Kurabayashi
总结: 端点监督连续风格滑块控制人体运动扩散。
方法: 风格嵌入方向加标量强度条件扩散与潜在正则。
证据: 无需中间强度真值实现平滑单调风格缩放。
为什么适合我: 支持人体动作风格重定向与扩散全身真机控制。
推荐理由: 人体动作扩散的风格强度滑条,偏动画风格编辑,不是可跟踪的运动先验或机器人控制。
原摘要

Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit, but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised continuous control. Given a content motion and a style motion, we construct a style direction in a learned motion-style embedding space and condition diffusion generation with a scalar intensity. The training objective combines diffusion denoising with latent intensity regularization to encourage smooth and monotonic style scaling without requiring intermediate-intensity ground-truth motions. Our framework is compatible with pretrained motion diffusion backbones and supports heterogeneous style datasets, including the multi-actor style motion dataset. To test out-of-range usability, we additionally introduce a small real-capture over-reaction extension and evaluate large-intensity behavior against these unseen targets. Experiments measure controllability, interpolation/extrapolation behavior, content preservation, and motion realism, with ablations on direction construction and loss design.

Towards VLA-Dreamer: Refining VLA Behavior Using World Models

2.0/5 偏低 裁判分 7.0 Strong 当日相对
Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter
总结: 在VLA嵌入空间训练世界模型精炼行为。
方法: 用视觉编码器嵌入预测未来动作相关动态。
证据: 概念验证嵌入可作非损失隐式世界模型接口。
为什么适合我: 提升VLA样本效率利于复杂地形感知运动智能。
推荐理由: VLA 嵌入空间世界模型的概念文,面向通用操作样本效率,不涉及人形全身或腿足运动。
原摘要

Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.

Luca Zanetti, Doganay Sirintuna, Idil Ozdamar, ... , Heng Zhang, Arash Ajoudani
总结: 自适应阻抗与ACT学习接触丰富操作演示。
方法: 遥操作自调运动方向刚度并记录用于ACT预测。
证据: 方向依赖刚度直接纳入演示无需手动选择。
为什么适合我: 结合模仿与阻抗支持接触操作合规全身控制。
推荐理由: 遥操作变阻抗加 ACT 的接触操作,平台是末端执行器而非全身 loco-manipulation。
原摘要

Contact-rich manipulation requires robots to balance accurate motion tracking with compliant interaction, yet most visual-action policies leave compliance fixed at the controller level. We present Imp-ACT, a methodologically grounded and practical approach to incorporating direction-dependent Cartesian stiffness modulation directly into demonstration collection, without manual stiffness selection or offline target reconstruction. During teleoperation, a self-tuning impedance controller adapts stiffness along the instantaneous direction of motion while maintaining compliance in orthogonal directions. The adapted stiffness is applied and recorded alongside visual observations and motion commands, capturing motion and compliance under the same dynamics. We implement this pipeline using Action Chunking with Transformer (ACT) to predict end-effector pose, gripper action, and motion-direction stiffness from visual, proprioceptive, and wrench observations. The performance of Imp-ACT is evaluated on wiping and plug insertion using both success rate and quantitative measures of contact behavior. Compared with fixed low- and high-stiffness baselines, Imp-ACT achieves comparable or higher success while maintaining low interaction forces. In wiping, it reduces contact-force vibration by approximately $29\times$ relative to the compliant baseline and $180\times$ relative to the stiff baseline. In plug insertion, it reduces forces orthogonal to the insertion direction by $43\%$ relative to the better fixed-stiffness baseline. These results highlight the benefit of maintaining sufficient stiffness along the direction needed for task execution while preserving compliance in other directions to limit contact forces and accommodate environmental constraints.