Papers for 2026-09-14

10 papers

CHOREO: Every Humanoid Skill as a Trajectory

5.0/5 很相关 裁判分 10.0 Top pick 当日相对
Ziyi Sun, Jingwen Chen, Yuxi Wang, ... , Zhaoxiang Zhang, Yujun Dong
总结: 提出训练无关的异构人形技能组合框架。
方法: 技能统一为SkillMotion,经延续混合或桥接组合。
证据: Unitree G1上2950资产,130序列成功率95.4%。
为什么适合我: 利于人形跑酷与loco-manipulation技能组合。
推荐理由: 直接对应人形全身控制与技能组合:把异构技能统一成带接触与边界条件的可执行轨迹,并在Unitree G1上做免训练衔接,紧贴全身运动跟踪与接触丰富控制。
原摘要

Recent advances in humanoid robotics have produced diverse skills through reinforcement learning, motion imitation, and generative modeling. Yet these capabilities remain siloed because they are built around incompatible representations, interfaces, and controllers. We present CHOREO, a framework for training-free composition of heterogeneous humanoid skills. Our key observation is that, regardless of how a skill is learned, it can ultimately be expressed as an executable motion trajectory. Based on this observation, CHOREO converts each capability into SkillMotion, a unified representation that combines motion states, contacts, semantics, and boundary conditions. Skills are composed through direct continuation, cubic Hermite blending, or validated bridge motions, without retraining source models or updating models at test time. On Unitree G1 in MuJoCo, CHOREO organizes 2,950 admitted SkillMotion assets derived from heterogeneous sources and achieves 95.4\% sequence success across 130 multi-action tasks, including 93.8\% success on eight-action sequences. These results demonstrate that executable trajectories provide a scalable interface for accumulating and composing pretrained humanoid capabilities.

Tsung-Chi Lin, Juo-Tung Chen, Chien-Ming Huang
总结: 探究全身遥操作在多样目标约束下的策略。
方法: 混合自由形式与约束控制用于TIAGo全身遥操作。
证据: 用户研究显示协调控制提升效率减关节极限风险。
为什么适合我: 启发全身感知运动与接触场景遥操作设计。
推荐理由: 最接近人形遥操作,但对象是TIAGo轮式移动操作臂的用户策略研究,不涉及人形/腿足运动跟踪、重定向或真机全身控制。
原摘要

This work investigates the control strategies of complex whole-body robot teleoperation that coordinate active perception, bimanual manipulation, and navigation. We developed a hybrid control framework, combining the free-form and constrained control, for the whole-body teleoperation of the TIAGo mobile manipulator. We conducted a user study to explore people's control strategies under different task constraints such as limited time and low tolerance of errors. Our results highlight the effective use of coordinated control in improving task efficiency and reducing the risk of reaching individual joint limits. We discuss our results and their implications for designing future whole-body robot teleoperation systems.

UniMo: Unifying Human and Animal Motion Generation

3.0/5 一般 裁判分 8.0 Strong 当日相对
Zeyu Zhang, Zhiyuan Zhang, Siheng Wang, ... , Ian Reid, Richard Hartley
总结: 统一人类动物运动生成的点云框架。
方法: 点云绕过拓扑差异,动态采样,配大规模数据集。
证据: 解决骨架多样与数据规模质量限制。
为什么适合我: 支持人体动作生成重定向至机器人控制。
推荐理由: 最接近人体动作与运动先验生成,但是跨物种点云动作生成,没有机器人跟踪、重定向或全身控制。
原摘要

The conditional generation of 3D motion has emerged as a key research topic due to its wide applicability across robotics, AR/VR, gaming, and content creation. However, extending recent advances in text-driven human motion generation to the animal domain remains challenging due to two core limitations. First, animals exhibit highly diverse skeletal topologies, unlike the standard human structure, making unified modeling across species difficult and leading to inefficient per-species models. Second, existing animal motion datasets suffer from limited scale and annotation quality, constraining model performance. To address these challenges, we propose UniMo, a unified point cloud-based motion generation framework that bypasses topological discrepancies by converting parametric skeletons into unparametric representations, further enhanced by dynamic sampling that allocates more points to active joints. Additionally, we present UniML3D, a large-scale motion-language dataset spanning both human and animal categories, containing 145,907 motion sequences and 433,388 captions-over 102x larger than existing animal datasets. Our method achieves state-of-the-art results on UniML3D and three public benchmarks including HumanML3D, KIT-ML, and AnimalML3D, demonstrating the feasibility and effectiveness of unified human-animal motion generation. Website: https://steve-zeyu-zhang.github.io/UniMo.

DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization

2.0/5 偏低 裁判分 9.0 Top pick 当日相对
Arjun Sohal, Yuchi Zhao, Miroslav Bogdanovic, Alan Aspuru-Guzik
总结: 扩散策略优化的去噪中间优势方法。
方法: 学习部分去噪动作价值构建去噪级优势。
证据: 改进扩散策略RL微调的信用分配。
为什么适合我: 打通扩散模型与强化学习用于策略优化。
推荐理由: 方法上靠近扩散策略与强化学习微调,但场景是通用操作而非人形全身控制、腿足感知运动或loco-manipulation。
原摘要

Diffusion-based robot policies have become widely used in robotic manipulation, where they are typically trained with behavior cloning. However, policies trained purely from demonstrations are limited by the quality and coverage of the available data. Reinforcement learning can further improve the performance of these pretrained policies through interaction. A common approach is to use policy-gradient methods that formulate diffusion-policy fine-tuning as an outer environment MDP together with an inner denoising MDP. However, existing methods typically assign the same environment-level credit to all denoising steps used to construct an action chunk, without distinguishing which intermediate decisions contributed most to the final return. We introduce Denoising Intermediate Advantage (DIA), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising level advantage for each step of the generative process. DIA combines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain. Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.

Zhenfeng Gan, Yanbo Chen, Lirong Che, ... , Rongkai Zhu, Xueqian Wang
总结: 语言引导地形自适应神经MPC穿越履带机器人。
方法: 神经运动学预测加NMPC,LLM安全更新参数。
证据: 全周期100ms内,三任务与多高度泛化验证。
为什么适合我: 复杂地形接触丰富场景的MPC感知运动相关。
推荐理由: 涉及地形自适应与神经MPC,但平台是铰接履带机器人,不是腿足/人形感知运动。
原摘要

In urban search and rescue, articulated tracked robots (ATRs) must traverse structured but contact-rich environments such as stairwells and cluttered building interiors. Reliable autonomy remains challenging because robot-terrain interaction (RTI) is hybrid and discontinuous, and effective flipper-track coordination is difficult to model analytically. We present ASTRIL-MPC, a language-guided neural kinematics model predictive control (MPC) framework for autonomous traversal. A learned kinematics model predicts short-horizon task-state increments from a height sequence and recent trajectories; NMPC plans with multi-objective costs and strict feasibility constraints; and a large language model (LLM) proposes bounded updates to selected weights and bounds through a safety-checked interface with range clipping, rate limiting, and consistency checks. The compiled predictor enables a full control cycle within 100 ms. Across three traversal tasks and a multi-height generalization setting, ASTRIL-MPC improves an aggregate traversal-quality score by up to 71% over a non-adaptive NMPC and by 67% over a PPO baseline, while eliminating measurable collision impacts during descent. These results indicate that combining terrain-conditioned neural kinematics, optimization-based planning, and language-guided adaptation yields data-efficient and robust autonomy for articulated tracked robots. Real-robot trials over four indoor obstacles further demonstrate transfer to contact-rich physical traversal.

Shogo Iwakata, Tomohiro Motoda, Ryosuke Yamada, ... , Shigeo Morishima, Yukiyasu Domae
总结: 几何先验预训练提升操作模仿学习效率。
方法: 自动生成平面物体手几何视觉预训练数据。
证据: 预训练暴露手物几何关系,无需任务特定数据。
为什么适合我: 模仿学习先验提升loco-manipulation样本效率。
推荐理由: 仅在模仿学习上弱相关,内容是桌面操作的几何预训练,不是全身运动或场景交互。
原摘要

Applying an imitation learning policy to a new manipulation task usually requires collecting new demonstrations and retraining the model, which makes sample efficiency a practical concern. Pretraining on large-scale robot datasets is effective in this respect, but such datasets are costly to collect and train on, while data augmentation techniques typically require a new round of data generation and retraining for each task. A complementary question is what useful prior can be provided to a policy at negligible cost before any task-specific data are collected. In this study, we construct a geometric visual pretraining dataset in which each scene contains only a plane, an object, and a hand, and trajectories are generated automatically. The scenes contain neither textures nor backgrounds; pretraining primarily exposes the policy to the geometric relationship between the hand and the object. Furthermore, representing the hand as a cube avoids tailoring the dataset to a specific robot morphology. We evaluate this geometric prior using ACT on three simulated robots across five manipulation tasks each, as well as on three real-world robot tasks. Across many of these robot--task combinations, fine-tuning from the geometric prior achieves higher success rates in the early stages of training than training from scratch while using only a small number of task demonstrations. These results suggest that even highly simplified geometric scenes can provide a useful initialization that transfers across robots and to real-world tasks when task data are limited.

Decentralized Evolution of Hexapod Gaits with Independent Leg Controllers

2.0/5 偏低 裁判分 4.0 Candidate 当日相对
Gary B. Parker, John Asaro, Jim O'Connor
总结: 六足独立腿控制器去中心化进化步态。
方法: 去中心化进化优化各腿,涌现协调行为。
证据: 优于合作协同进化,生成稳定自适应步态。
为什么适合我: 腿足机器人步态与全身运动智能相关。
推荐理由: 属于腿足步态,但是六足分散进化控制,不涉及人形全身智能、模仿或感知跑酷。
原摘要

This paper presents a novel approach to hexapod locomotion by evolving each leg's gait independently through a decentralized evolutionary algorithm. Using the Webots simulator and the Mantis hexapod robot, we optimize individual leg controllers without centralized coordination, allowing emergent behaviors to drive the development of efficient, coordinated locomotion. Our decentralized method is benchmarked against cooperative coevolution, demonstrating improved efficacy in generating stable and adaptive gaits while showing interesting emergent coordination. By enabling independent evolution of leg controllers, this method reduces the complexity of gait optimization and highlights the potential of decentralized strategies for scalable and adaptive robotic systems.

Ryunosuke Yamada, Tomohiro Motoda, Yukiyasu Domae, Tokuo Tsuji
总结: 在线材料估计条件扩散策略塑造DLO。
方法: 循环网估计材料标签条件扩散策略每步。
证据: 480演示,真值材料条件提升平均成功率。
为什么适合我: 扩散策略用于接触丰富可变形操作。
推荐理由: 可变形线状物体的材料条件扩散策略,与人形、腿足运动和loco-manipulation无关。
原摘要

Shape control of deformable linear objects (DLOs) is challenging for imitation learning because deformation behavior varies with material properties such as stiffness and elasticity, so a single policy must generate different action sequences for different objects even when the goal shape is identical. We propose a diffusion policy conditioned on material labels that are estimated online during manipulation. A recurrent estimation network predicts the material label of the grasped object from the time series of multi-view images and robot joint states, and the predicted label conditions the diffusion policy at every inference step. We collected 480 real-robot demonstrations covering four DLO materials and three groove-placement tasks, and compared per-material specialist policies, a task-conditioned policy without material labels, a policy conditioned on ground-truth material labels, and the proposed policy. Conditioning on ground-truth material labels improved the average success rate from 45.8% to 60.0% over the task-only policy, and the proposed policy reached 60.8% without any prior material information, matching the policy given ground-truth labels. A post-hoc analysis shows that the estimator extracts material-related information from the manipulation observations and that the diffusion policy responds to the resulting conditioning signal, while the one pronounced failure case is associated with persistent confusion between two similar materials.

Lennart Werner, Pol Eyschen, Sean Costello, ... , Andrei Cramariuc, Marco Hutter
总结: 材料状态RL实现挖掘机可迁移土壤操作。
方法: MPM仿真RL条件材料状态,归一化末端部署。
证据: 11.5t挖掘机与桌面机验证,建42m结构。
为什么适合我: 启发接触丰富可变形场景的强化学习。
推荐理由: 挖掘机土壤操作的材料状态强化学习,与人形/腿足全身运动无关。
原摘要

Earthmoving tasks such as excavation, backfilling, or embankment construction require deliberate repositioning of deformable soil. For these tasks, human operators use all shovel faces, while autonomous systems so far are limited to excavation and dumping. Current methods often rely on heuristic models but do not incorporate soil mechanics. We address this shortcoming by using Reinforcement Learning in a GPU-parallelized Material Point Method particle simulation. Our controllers are conditioned on material state such as shape and compactness, enabling skills that use multiple contact faces of the tool and displace material both inside and outside of the shovel. To use the same learned weights across machines, our policies operate in a normalized end-effector space and are deployed through a calibrated machine interface. We evaluate this calibrated transfer on an 11.5t hydraulic excavator and a 500g tabletop robot. We validate performance through autonomous construction of a 42m long, 2.1m high embankment in 45min, executing 201 individual policy strokes without failure, retry, or operator intervention. In a direct comparison, the autonomous controller matches an expert operator's progression speed and produces a higher, more consistent embankment. Additional qualitative backfilling and compaction experiments demonstrate the material-state awareness and calibrated transfer across machines.

SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy

1.0/5 偏低 裁判分 3.0 Candidate 当日相对
Xiefeng Wu, Shu Zhang, Zhaojie Chu, Mingyu Hu
总结: sigmoid有界熵稳定保守Q学习。
方法: 替换log熵为严格正sigmoid有界熵。
证据: D4RL与视觉任务匹配超基线,训练更稳。
为什么适合我: 稳定离线在线RL利于真机全身控制。
推荐理由: 通用离线到在线Q学习稳定性方法,没有人形或腿足运动问题。
原摘要

Offline-to-online reinforcement learning reduces interaction cost for real-world robot learning but suffers from persistent value estimation instability. Existing methods address this through pessimistic regularization, lower-bound calibration, and architectural normalization, but an overlooked source of instability lies in the entropy formulation: the standard log-entropy term can become negative, destabilizing policy updates. We introduce SCQ (Sigmoid-Bounded Conservative Q-Learning), which replaces this term with a sigmoid-bounded formulation that stays strictly positive. SCQ retains conservative Q regularization and return-based lower-bound calibration, stabilizing policy optimization without sacrificing exploration. We evaluate SCQ on D4RL (Minari) benchmarks under both single-demonstration and standard dataset settings, as well as on simulation and real-world visual tasks. SCQ matches or exceeds baseline performance while exhibiting more stable training dynamics across state-based and visual benchmarks, and transfers to four real-robot platforms including manipulation, wheeled, quadruped, and humanoid systems. A direct clipping intervention that removes negative log-probability contributions, together with gradient-matched positive-score controls, indicates that positivity rather than a particular score shape alone drives much of the improvement. Project website: https://scq-rl.github.io.