Papers for 2026-09-21

10 papers
Jiacheng Wei, Jerry Bai, Xiaoyu Yue, ... , Chenyi Chen, Yixiao Ge
总结: 统一世界模型联合预测动作与未来视频并作模拟器。
方法: 共享视频骨干用无标视频学动态有标演示联合预测粗到细课程。
证据: 连接异构经验学动作并生成新经验改进策略超越机器人演示。
为什么适合我: 异构经验与视频预测利于全身运动模仿与可迁移控制器学习。
原摘要

A general-purpose robot needs to draw on diverse experience, choose actions, and anticipate how those actions will change the world. We introduce XPACE, a unified embodied world model that serves as both a world action model, jointly predicting executable robot actions and future video, and a world simulator, predicting the visual consequences of prescribed actions. Our key insight is that video prediction can both connect heterogeneous experience to action learning and generate new experience for policy improvement. With a shared video backbone between the policy and simulator, we use action-unlabeled video to learn visual dynamics and action-labeled human and robot demonstrations to jointly learn video and action prediction. Building on this architecture, a coarse-to-fine training curriculum progressively emphasizes robot control while retaining human experience, allowing the policy to learn behaviors beyond those covered by robot demonstrations. Beyond learning from recorded experience, XPACE uses its simulator to create additional recovery supervision for the policy. Specifically, we adapt the simulator to its own generated context, synthesize deviation-recovery trajectories around expert demonstrations, and fine-tune the policy on filtered recovery examples. Experiments on XPENG's IRON humanoid robot show that heterogeneous training improves robustness and enables transfer of human-observed skills to tasks absent from robot demonstrations, while recovery data generated by the model's own simulator further improves real-world task completion. Together, these results demonstrate how joint world and action modeling connects learning from heterogeneous experience with simulation-driven policy self-improvement.

Xiangyu Li, Tiancheng Lai, Xijie Huang, ... , Fei Gao, Yanjun Cao
总结: 零样本框架用视频想象实现陆空机器人语言引导导航。
方法: VLM生成提示选视频重建带模式航点两阶段校准尺度几何。
证据: 无需任务训练从视频恢复度量一致导航参考。
为什么适合我: 视频想象与几何校准启发腿式感知运动与地形接触利用。
原摘要

Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.

Xiangrui Wang, Yuetong Wu, Jalen Beeman, ... , Read Hayes, Boyi Hu
总结: 开放多模态数据集记录零售中人类全身活动与机器人引导。
方法: 十六人三条件货架拣选用惯性与肌电传感器记录运动。
证据: 展示不同机器人交互下全身运动强度与肌肉激活变化。
为什么适合我: 人类全身运动数据支持运动重定向与模仿全身控制器。
原摘要

Open datasets describing human behavior in environments shared with mobile robots remain limited, particularly for retail activities that combine locomotion, reaching, object handling, and robot guided movement. This paper introduces OHRID Retail, an open, human centered multimodal dataset collected from 16 healthy adults performing a simulated shelf picking task under three within participant conditions: no robot, low speed robot guidance, and high speed robot guidance. Each participant completed two trials per condition. Whole body kinematics were recorded using 17 Xsens Awinda inertial sensors and muscle activity was measured at 10 locations using Delsys Trigno surface electromyography sensors. Descriptive analyses demonstrate variation in whole body movement intensity and muscle activation across robot interaction conditions and body locations. OHRID Retail provides openly available raw recordings, processed measures, documentation, and reproducible analysis resources. The dataset can support research in human activity recognition, multimodal sensor fusion, occupational biomechanics, ergonomics, human aware robot navigation, and human robot interaction in retail and related shared environments.

Masato Kobayashi, Takeru Tsunoori
总结: 水下双臂机器人学习平台支持遥操作演示与策略部署。
方法: 自定义主从硬件扩展LeRobot集成传感收集训练与部署。
证据: 演示臂间传递共享操作等并评估多种策略水下表现。
为什么适合我: 双臂遥操作与模仿可延伸人形移动操作与全身控制。
原摘要

Underwater visuomotor policy learning has focused primarily on single manipulators, while bimanual imitation learning has been studied largely in air. We present ULOHA, an underwater bimanual robot learning platform that combines custom-designed leader--follower hardware with software extensions to LeRobot, integrating teleoperation, multi-view sensing, demonstration collection, policy training, and autonomous deployment. Real-robot experiments demonstrate a range of coordinated underwater bimanual behaviors, including inter-arm transfer, shared-object manipulation, and buoyancy-driven interception. We evaluate ACT, Diffusion Policy, and the vision--language--action model SmolVLA on the platform. We investigate how learning methods and execution strategies developed for manipulation in air perform underwater, examining bubble disturbances, buoyancy-driven object motion, action-execution horizons, and real-time chunking. A separate single-arm study examines policy transfer between air and water and shows that demonstrations spanning both media support execution in both under the tested conditions. ULOHA provides a unified experimental platform for studying underwater bimanual robot learning under the coupled perceptual and physical effects of underwater environments. Additional material: https://mertcookimg.github.io/uloha/

Ruben Beumer, Tom Janssen, René van de Molengraft, Duarte Antunes
总结: 四足机器人装铣刀工具实现机械精准除草减土壤压实。
方法: Spot刚性装工具用腿自由度驱动脚静止自主除草架构。
证据: 分析原型精度与效率并集成自主除草系统。
为什么适合我: 四足接触丰富操作利用腿自由度契合全身运动控制。
原摘要

Herbicide-based weed control is increasingly unsustainable due to rising weed resistance and the adverse environmental impacts of chemical use. While mechanical weed control avoids these drawbacks, it is typically implemented using large machines that cause soil compaction. We propose a novel alternative based on small mobile robots for mechanical weeding. Compared with existing automated mechanical weeding approaches, the proposed method offers reduced soil compaction, simpler automation, and improved scalability. Our solution involves a Boston Dynamics Spot quadruped robot equipped with a custom weed removal tool featuring a milling bit at its end. The tool is rigidly attached to the robot and uses the degrees of freedom of the robot base by actuating the legs, while keeping the feet stationary. We develop a software architecture that enables autonomous weed removal and integrate this system with all other required components. We analyze the accuracy and efficiency of the current proof of concept both in an indoor and outdoor environment and provide recommendations for future work to make the system more accurate and efficient.

Toshiki Otani, Hiromu Taketsugu, Norimichi Ukita
总结: 能量正则模仿学习实现力与功感知机器人操作。
方法: 定义关节功代理训可微预测器作正则微调预训练策略。
证据: 十二任务平均功从208.8J降至204.4J成功率略升。
为什么适合我: 功感知正则可提升接触丰富全身运动控制物理效率。
原摘要

This paper studies energy-aware manipulation as a physically grounded learning problem. We define a joint-space mechanical-work proxy from joint torque and angular displacement, and train a differentiable energy predictor that estimates this work from robot states and actions. The predictor converts a non-differentiable simulator-side physical quantity into a differentiable regularizer for fine-tuning a pretrained manipulation policy. We instantiate the framework with RVT-2 on RLBench and evaluate 12 manipulation tasks involving object contact, articulated motion, placement, pushing, and sweeping. The proposed fine-tuning reduces the average mechanical work from 208.8J to 204.4J (i.e., 2.1% reduction), while the mean task success rate also increases slightly from 86.2% to 86.9%. These results show that work-aware policy optimization can suppress physically inefficient motion without requiring an explicit differentiable dynamics model.

Pengqin Wang, Qiming Zhang, Shaojie Shen, Jun Ma
总结: 无评论家强化学习框架用于连续机器人控制。
方法: 并行轨迹估折扣回报组归一化用相对优势裁剪更新策略。
证据: 实例化框架评估连续控制有效性。
为什么适合我: 简化RL利于腿式敏捷行走与全身控制器策略优化。
原摘要

Actor-critic architecture has been widely used in continuous robot control. However, they rely on learning a value network, introducing additional computational overhead during training. Moreover, policy learning may also be affected by the approximation error of value estimation. Critic-free group relative policy optimization methods provide a simpler training approach by removing the need for a critic. However, they fail to learn long-term action outcomes when directly applying immediate rewards to policy optimization in dense-reward environments. To address these problems, we propose Group Relative Return Policy Optimization (GR2PO), a critic-free reinforcement learning framework for continuous robot control. GR2PO estimates the discounted returns from the parallelly collected trajectories, performs group normalization at each rollout time index, and uses relative advantages and clipped targets to update the policy. To evaluate the effectiveness of the proposed framework, we instantiate it on robot control simulation environments and deploy the model to a real-world edge device. The results show that GR2PO significantly outperforms critic-free baselines that use immediate rewards and performs competitively against state-of-the-art actor-critic methods. Furthermore, GR2PO demonstrates competitive training efficiency. Inference tests on NVIDIA Jetson TX2 demonstrate the feasibility of deploying the learned policies on edge platforms. Further ablation experiments analyze the effects of parallel group size, return estimation methods, and target clipping ratio on learning performance. To support follow-up research, we will make the complete code publicly available after the paper is accepted, including the framework implementation, experimental configuration, and training and evaluation scripts.

Mehdi Heydari Shahna, Seihun Kim, Soyi Jung, ... , Jouni Mattila, Joongheon Kim
总结: 模仿学习加自适应笛卡尔跟踪控制三十吨挖掘机。
方法: 模仿策略生名义速率自适应反馈校正后停止距离约束。
证据: 仿真完成所有目标时长更短终端误差更低。
为什么适合我: 自适应跟踪启发人形全身运动跟踪与遥操作控制。
原摘要

Autonomous excavator control is challenged by coupled kinematics, actuation lag, and uncertainty. We propose imitation learning and adaptive Cartesian tracking (IL-ACT), a novel motion control framework for a 30-ton-class excavator. An anchored, 14-input imitation policy pretrained on operator demonstrations generates nominal joint rates; adaptive Cartesian feedback and gated gain/bias estimation correct these commands before a stopping-distance governor constrains joint-reference generation. Simscape evaluation covers 100 sequential goals and spiral, figure-eight, and rounded-raster tracking, including 88 additional runs across three training seeds, two initializations, and speeds, under hydraulic response and sensing conditions. Compared with Teacher+ACT, IL-ACT completes all goals with shorter duration and lower terminal errors under both response conditions. Telemetry-initialized IL-ACT lowers RMSE in all 24 figure-eight and rounded-raster seed comparisons and lowers additional-load spiral mean RMSE by approximately 29%. Original spiral RMSE also improves over IL-only and PID. Under a shared sensor-noise realization, telemetry-initialized IL-ACT achieves 27.67% lower mean RMSE than Teacher+ACT; enabling estimation reduces mean RMSE by $22.44\%$ relative to the frozen estimator. Pretrained-weight effects remain mixed, and the original teacher comparison exhibits a spiral RMSE--maximum-error tradeoff. Analysis establishes bounded adaptive states and Cartesian feedback, with reference admissibility conditional on governor feasibility.

Sun Woo Kim, Xue Bin Peng
总结: 扩散技能发现学习多样可重用运动技能。
方法: 最大化技能潜变量与状态互信息用扩散促进多样覆盖。
证据: 技能库支持广泛下游任务含空间时间变化。
为什么适合我: 多样可重用技能直接服务仿真角色模仿与全身控制。
原摘要

Humans efficiently learn new tasks by reusing a rich repertoire of motor skills across different goals and contexts. A similar strategy can also be used to enable simulated characters to efficiently perform new tasks by leveraging reusable motor skills. To support a wide range of downstream tasks, the learned repertoire should be diverse, consisting of distinct behaviors as well as spatial and temporal variation within each behavior. A commonly used method for learning diverse skills is by maximizing the mutual information between skill latents and the states produced by a policy. The marginal state entropy promotes broad behavioral coverage, while the conditional entropy encourages consistent behaviors from each latent. However, directly estimating the marginal state entropy is intractable in high-dimensional control problems. Prior methods therefore rely on indirect latent-space approximations or coarse estimators of the state distribution. These approximations may not effectively promote broad coverage of the state space, resulting in skills with limited behavioral diversity and reduced utility for downstream tasks. In this work, we propose Diffusion Skill Discovery (DSD), a skill discovery method that uses a diffusion model to approximate the entropy gradient of the policy-induced state distribution through score matching. The resulting objective encourages the discovery of skills that produce a broader range of behaviors for high-dimensional humanoid control. The learned skills are reused in two downstream control settings: hierarchical control with a task-specific high-level policy and zero-shot control through latent selection from offline trajectories. Our experiments show that DSD discovers a broader repertoire of reusable motor skills than prior skill discovery methods, leading to the emergence of complex and agile behaviors that can be reused across downstream tasks.

Xingpeng Sun, Zherong Pan, Kai Cheng, ... , Syed Talha Bukhari, Aniket Bera
总结: RL微调LLM生成MIP代码解决多模态运动规划。
方法: LLM分解任务为MIP变量约束目标生成可执行代码。
证据: 处理模式转换与连续动态如双足行走后抓取。
为什么适合我: 多模态规划含双足操作契合人形全身运动与接触。
原摘要

Multi-Modal Motion Planning (M$^3$P) requires joint reasoning over continuous motions and discrete mode transitions, making it difficult to solve efficiently. For instance, a bipedal robot may walk to a target location and then use its arms to grasp an object. This scenario captures both mode transitions and continuous dynamics, yielding feasible paths that neither purely discrete nor continuous planners can handle. While Mixed-Integer Programming (MIP) offers a principled framework, constructing tractable formulations for non-convex problems is typically manual and domain-specific, especially in the approximate, discretization-based MIP regime needed for non-convex robotic tasks. We propose M$^3$P-R1, a reinforcement learning method that fine-tunes large language models (LLMs) to decompose M$^3$P tasks into MIP variables, constraints, and objectives. Instead of directly outputting answers, which are often prone to hallucination, the model generates executable Python code using MIP optimization libraries and constraint interfaces. This enables solver-backed execution for robust and verifiable solutions. Trained with an outcome-driven reward against the solver, M$^3$P-R1 learns to compose modality-level discretization primitives and synthesize cross-modal coupling constraints, producing executable MIP programs for complex M$^3$P tasks.