Papers for 2026-10-04

10 papers
Huimin Pan, Yufan Ren, Kunpeng Song, ... , Xiaoyang Guo, Chenyi Chen
总结: 通过以自我为中心的人类视频预训练VLA模型,用相机空间动作表示绕过本体差异与人手重定向,实现人形灵巧操作的规模化。
方法: 将动作定义在自我视频原生的相机空间并对齐人机动作维度语义,融合异构机器人数据开展250至10000小时的规模化预训练。
证据: 验证损失随250到10000小时总预训练预算近似对数线性下降,更大预算带来更强的下游操作策略性能。
为什么适合我: 绕过显体重定向的跨本体动作表征思路,可启发我摆脱参考动作对精确身体运动学依赖的全身跟踪与数据扩展。
原摘要

Egocentric human video offers a scalable data source for dexterous manipulation, yet using it to train humanoid robots presents two challenges: (1) an embodiment gap, as human hands differ structurally from robot end-effectors and low-cost egocentric recordings lack the torso kinematics required by conventional retargeting; and (2) heterogeneous data quality, including noisy hand-pose tracking and weakly aligned text annotations. We introduce IronMind, a vision-language-action (VLA) model that uses egocentric human video and heterogeneous robot data to pretrain policies for humanoid dexterous manipulation. To bridge the embodiment gap, IronMind bypasses explicit body-retargeting by using a camera-space action representation, the native reference space of egocentric video, and semantically aligning robot and human action dimensions. Across total pretraining budgets from 250 to 10,000 hours, validation loss decreases approximately log-linearly with data scale. Larger pretraining budgets also improve out-of-distribution real-robot manipulation after post-training: across six challenging tasks with unseen objects, affordances, and reasoning prompts, the 10,000-hour model achieves a 55.0% success rate, compared with at most 11.7% for every pretraining budget up to 5,000 hours and 5.0% without pretraining. At the same pretraining budget, the camera-space action representation also outperforms the torso-frame baseline. Together, these findings support pretraining with a camera-space action representation on large-scale human egocentric data as a scalable foundation for humanoid robot manipulation.

Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman
总结: 用LLM合成参数化反应式控制器,结构由大模型编写、参数靠免导数搜索拟合,决策时无需LLM推理或规划即可实时执行。
方法: 分离程序结构与参数:LLM生成Python控制器骨架,免导数搜索利用环境反馈拟合连续控制所需的参数。
证据: 在Atari、Flappy Bird与MuJoCo上超越基于规划的程序合成,动作选择快于PPO且环境交互更少,可跨任务变体迁移。
为什么适合我: 免推理的参数化反馈控制器可作为学习型全身策略的实时补充或后备层,有利于真实机器人低延迟安全部署。
原摘要

Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.

Parastoo Ali Pour, David R. Martin, Chang Min Hur, ... , Pramod Khargonekar, Mohammad Abdullah Al Faruque
总结: 在Unitree G1上验证单操作员遥操作完成建筑任务的可行性,XR上肢控制与脚踏行走并行实现边走边操作。
方法: 结合扩展现实的臂手遥操作与脚踏式行走控制,支持同时操纵与移动,并在O*NET选出的两个建筑任务上评测。
证据: 工具搬运成功率100%、表面喷涂80%,任务均能完成,但遥操作耗时显著长于人工执行基线。
为什么适合我: 同时行走与操作的遥操作系统既能短期落地人形作业,又能批量采集我所需的高质量全身移动操作演示数据。
原摘要

We present a teleoperation system that enables a single operator to perform construction tasks on a Unitree G1 humanoid, combining extended reality (XR) based upper body control with pedal-based locomotion to enable simultaneous manipulation and locomotion. Motivated by persistent labor shortages, hazardous working conditions, and challenges in humanoid autonomy, we investigate teleoperation as a practical near-term approach for reducing physical strain on workers while generating high quality demonstration data. We evaluate the system on two representative construction tasks drawn from O*NET occupational database, and report task success and completion time relative to a manual baseline. The system achieved 100% success on tool transport and 80% success on surface painting, with teleoperation requiring substantially more time compared to manual execution.

Merkourios Simos, Chengkun Li, Bianca Ziliotto, Alexander Mathis
总结: 从纯运动学轨迹恢复地形支撑几何并重定向到肌肉骨骼模型,训练出单一的地形感知肌肉驱动行走策略。
方法: 融合地形先验、接触估计与负自由空间证据重建支撑几何,重定向时施加解剖、肌腱连续性与接触一致性约束。
证据: 用五个数据集导出的9.4小时运动-地形对训练单一策略,在重建、重定向与留存跟踪基准上均表现良好。
为什么适合我: 从无地形标注的运动数据恢复接触几何的管线,可直接增强我的人形控制器对地形支撑与接触信息的利用。
原摘要

Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human motions. Yet these capabilities remain largely confined to flat ground, in part because motion datasets rarely include aligned terrain geometry and because retargeting terrain interactions to complex musculoskeletal bodies is challenging. We present TERRA, an end-to-end pipeline for terrain-aware retargeting and control of musculoskeletal locomotion. From kinematic trajectories alone, TERRA combines terrain priors, estimated contacts, and negative free-space evidence to recover task-relevant support geometry. TERRA further considers anatomical, tendon-continuity, and contact constraints during retargeting. Using the resulting motion-terrain pairs from five datasets, we successfully train a single muscle-actuated control policy on 9.4 hours of diverse locomotion. Across reconstruction, retargeting, and held-out tracking benchmarks, TERRA improves terrain accuracy, sharply reduces anatomical and interaction violations, and achieves the highest observed completion rate over supported terrain families. Overall, TERRA provides a practical route from scene-less motion data to muscle-actuated locomotion over diverse non-flat terrain. Project website: https://cnai.epfl.ch/terra/

Galbot Team, Xuchuan Chen, Xiaoqian Cheng, ... , Yixin Zheng, Weiyi Zhu
总结: 系统评测GPT-6 Astra直接生成数值机器人动作的能力,发现导航与混合控制可用,灵巧手直接控制仍薄弱。
方法: 在六个域检验直接控制、与学习策略协作及反馈适应三种模式,覆盖抓取、灵巧、移动操作、导航与运动控制。
证据: 与π0.5混合控制获48%成功率,DexJoCo灵巧操作50%、RoboCasa365为38.7%,RxR导航达92%但搜索绕路明显。
为什么适合我: 结果印证通用大模型仍难协调精细接触,我关注的接触丰富全身控制仍需专门训练的学习型策略来支撑。
原摘要

GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. In gripper manipulation, Astra can correct task targets and prepare contact conditions for subsequent policy execution; hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset. In dexterous manipulation, hybrid control achieves 50% success in ten experience-guided DexJoCo trials, while direct in-hand control struggles to coordinate finger contacts. In mobile manipulation, hybrid control reaches 38.7% success on the evaluated RoboCasa365. In navigation, Astra leads our local comparisons, reaching 92% success on RxR instruction following and 82% on HM3D object search, although search incurs substantial detours. In locomotion, dense motion-reference generation remains unreliable: none of five sequential attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. In humanoid loco-manipulation, Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks with pretrained whole-body controllers. These findings reveal a gap between useful task decisions and reliable physical control. Inference latency further constrains practical control: across 50 RoboDojo instances per condition, policy-assisted and direct control consume 624.8 million and 1.132 billion tokens. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference.

Yizhao Li, Pusen Gao, Ming Wang, ... , Shuo Yang, Hao Xu
总结: 提出ECHO-G,以语音音频与时间对齐文本联合条件,直接在人形机器人动作空间生成说话伴随的全身动作。
方法: SGDiT扩散Transformer融合帧级声学与词元级语言特征,以整流流匹配在机器人空间训练,并构建音频-文本-机器人数据集与基准。
证据: 对比显示直接机器人空间生成优于人类动作生成加重定向管线,消融验证音文联合条件的收益,并完成实机部署。
为什么适合我: 生成式策略直接在机器人空间产出多模态条件全身动作,为我的人形表达性全身控制与部署提供了可行路径。
原摘要

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

Jingwei Jia, Keyu Zhou, Jiewei Wang, ... , Jin Wang, Shunlei Li
总结: 面向长时程手术辅助,提出人-人形交互规划框架RoboAssist,在线协调人类工作流与人形执行并全程保障安全。
方法: 非对称双轨表征分离部分可观测的人类过程状态与可执行任务序列,变化时仅重规划受影响后缀,叠加跨层安全架构。
证据: 在Unitree G1上演示了工作流推理、任务协调、近距交接调节与独立全身运行时监督的完整长时程辅助流程。
为什么适合我: 其独立的全身运行时安全监督与局部后缀重规划思想,可迁移到我的人形全身控制器在真实交互场景中的部署。
原摘要

Long-horizon surgical assistance requires humanoid robots to coordinate with evolving human activities while maintaining safety across planning and execution. We present RoboAssist, an agent-based framework for interactive human-humanoid planning that integrates workflow reasoning, task coordination, and cross-layer safety. At its core is an asymmetric dual-track representation that separates partially observed human process states from executable robot task sequences. By updating human-process estimates, scene context, and task dependencies online, RoboAssist revalidates the remaining task sequence and replans only the affected suffix when workflow requests change. A cross-layer safety architecture combines preventive navigation regulation, reactive regulation during close-range handover, and independent whole-body runtime supervision. This design couples online task coordination with safety constraints throughout execution. We demonstrate the framework on a Unitree G1 humanoid robot in long-horizon, multi-stage simulated surgical assistance scenarios encompassing multimodal interaction, instrument handling, medical material transport, navigation, and safe human-robot handover. Experiments show multi-stage task completion and adaptation to workflow-request changes. A targeted full-replanning ablation shows that residual replanning reduces plan-update latency and post-update token usage. Separate safety experiments demonstrate complementary protection across navigation, handover, and runtime supervision. Additional results and demonstrations are available online at https://roboassist.github.io.

Bosong Ding, Xianglin Zhang, Miao Xin, Murat Kirtay, Giacomo Spigler
总结: 提出GestAdapt,以规定的手腕工作空间为条件生成伴随语音手势,让机器人动作主动适应墙边等空间约束。
方法: 以共享动作表征从六个伴随语音语料库联合学习,按指定手腕工作空间条件化生成,并支持跨机器人本体重定向。
证据: 生成动作在尊重工作空间的同时贴近真实分布,用户研究质量均分3.24/5,高于无工作空间约束的基线。
为什么适合我: 把几何工作空间约束显式注入生成式动作策略,可启发我在全身生成与跟踪中融入地形与接触边界条件。
原摘要

Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. Since the same speech can be accompanied by different gestures, a robot can respond to workspace constraints, e.g., gestures for speech next to a wall. In these scenarios, the robot should gesture in a suitable motion rather than simply correcting an unconstrained one. To achieve this goal, we present GestAdapt, a workspace-conditioned framework that conditions co-speech gesture generation on a prescribed wrist workspace. The GestAdapt framework learns from six complementary co-speech corpora through a shared motion representation and supports retargeting to different robot embodiments. Quantitative evaluation shows that generated motions remain close to the real-motion distribution while respecting the workspace. In a user study, gestures generated under modified workspace constraints receive a mean quality score of 3.24/5, above our no-workspace variant (2.43/5) and below the reference motions (3.68/5). In a real robot evaluation, all compared motions are retargeted to the Reachy2 humanoid robot under identical workspace constraints. Motions generated with our framework rank first in 69.7\% of comparisons, higher than our no-workspace variant baseline and retargeted ground-truth motions constrained afterward. Overall, the results support adapting gestures to the available workspace during generation, rather than modifying unconstrained trajectories afterward to satisfy workspace constraints, potentially compromising gesture naturalness.

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini
总结: 用四足行走经验持续学习地形可通过性,从触前图像预测打滑、足载与能耗等尚未发生的接触交互结果。
方法: 冻结DINOv3提取触前视觉描述子,证据回归器加权映射到五个本体指标并输出偶然与认知不确定性,以回放加验证门控持续更新。
证据: 预测平面足端打滑、法向地反力、牵引指数、运输代价与触地加载率五项指标,并随新接触在线持续改进模型。
为什么适合我: 视觉到接触结果的预测正契合我的感知运动与地形利用目标,其持续学习与不确定性机制可直接借鉴。
原摘要

Safe and efficient quadruped navigation over unfamiliar terrain requires predicting terrain-robot interaction before contact: geometry and visual appearance alone cannot reveal how the robot will slip, load its feet, or expend energy. This paper presents a continual learning pipeline that uses locomotion experience to learn these interaction outcomes from pre-contact images and continually updates the predictions as new contacts are observed. Pre-contact descriptors, produced by a DINOv3 backbone model frozen during training, are mapped to five proprioceptive indicators weighted according to measurement reliability: planar foot slip, mean normal ground-reaction force, traction index, cost of transport, and touchdown loading rate. A compact evidential regressor allows us to predict these indicators together with aleatoric and epistemic uncertainty from the visual descriptors. Continual adaptation combines bounded experience replay with a validation gate: candidate models replace the deployed predictor only when they improve performance on recent held-out data while keeping degradation on historical held-out data within a prescribed tolerance. Predictions and epistemic uncertainty are projected into a local multilayer map and combined into a conservative traversability score map whose property weights can be adjusted without retraining. The resulting map is used for downstream navigation tests. The ROS2 implementation supports evaluation on a Unitree Go2 in simulation and on hardware, with models trained separately in each domain. On a sequential hardware stream over three previously unseen terrains, gated replay reduces final anchor negative log-likelihood (NLL) degradation by 23.1% relative to replay without the gate while attaining similar new-terrain adaptation.

Kyungmin Lee, Sibeen Kim, Dongyoon Hwang, ... , Jaegul Choo, Hojoon Lee
总结: 提出FlashDexRetarget,用强化学习把人手-物体演示高成功率重定向到灵巧手,加速多动作操作数据生成。
方法: 组合物体点云、手-物距离与未来轨迹编码及互补奖励,左右手分离actor-critic并将FlashSAC改造用于动作跟踪。
证据: 在50个单双物体交互动作基准上达约90%成功率,重定向成功率与训练效率均超过既有物理重定向方法。
为什么适合我: 其RL动作重定向与手-物关系跟踪可推广至全身运动重定向,帮助我批量构建可跟踪的参考动作数据集。
原摘要

Human hand-object demonstrations offer a reusable source of dexterous robot manipulation data, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based approaches face limitations in retargeting success, motion-specific training efficiency, or both. To address these limitations, we introduce FlashDexRetarget, an RL-based framework for high-success, efficient dexterous motion retargeting. To make the demonstrated interaction easier to learn, we combine object point-cloud observations, hand-object distance features, and future trajectory encodings with complementary rewards that supervise object motion and reference hand-object relationships. To further accelerate learning, we employ separate left- and right-hand actor critic networks and adapt the off-policy algorithm, FlashSAC to dexterous motion tracking. On a benchmark of 50 motions spanning single-object and two-object interactions, FlashDexRetarget achieves a 90% success rate, approximately 2.5x that of the evaluated sampling-based baselines, while requiring up to 100x less training compute than the evaluated RL-based baselines. Evaluations on both XHand and Sharpa Wave Hand show consistent gains, and component-wise ablations examine the contributions of our design choices. Beyond the 50-motion benchmark, experiments with 200, 500, and 1,000 motions demonstrate that our method remains stable at larger scales and produces successful retargeted motions more efficiently as the training set grows. Qualitative replay results using real-world-captured demonstrations further illustrate the applicability of our framework to recorded human manipulation. Videos and code are available at https://davian-robotics.github.io/FlashDexRetarget/