Papers for 2026-09-12

10 papers
Chaelin Kim, Seokhyeon Hong, Kwan Yun, ... , Inseo Jang, Junyong Noh
总结: 接触引导HOI运动重定向到多样人形角色。
方法: 部位感知运动嵌入自编码器编码全身运动。
证据: 保持运动语义与接触一致的交互模式。
为什么适合我: 支持人形loco-manipulation与人体动作重定向。
推荐理由: 直接对应人体动作重定向与模仿:接触引导地把人-物交互迁移到不同骨架的人形角色,并保持全身语义与接触一致,贴近全身 loco-manipulation。
原摘要

We present ReCHOIR, a novel contact-guided motion retargeting method for transferring human object interaction (HOI) motions across diverse humanoid characters. Unlike prior motion retargeting methods that primarily focus on transferring human motion alone, our goal is to preserve not only the semantics of the original body movement but also consistent interaction between the character and the manipulated object, while jointly producing aligned target human and object motions. Given source HOI motion, object geometry, and contact cues extracted from the source interaction, ReCHOIR retargets an HOI sequence to target characters with different skeletal configurations while maintaining both motion semantics and contact-consistent interaction patterns. Our method builds on a Part-Aware Motion Embedding (PAME) autoencoder, which encodes full-body motion into a shared body-part-wise latent space. This representation enables generalization across heterogeneous skeletons while preserving local motion semantics beneficial for part-aware adaptation in HOI retargeting. On top of this representation, we introduce a contact-guided retargeting module and an object motion decoder for HOI retargeting. The contact-guided retargeting module treats the source object interaction as a condition for refining target character motion: object- and contact-related signals are encoded into a body-part-aligned latent representation and injected into decoding through a residual control branch, enabling stronger adaptation in interaction-relevant body regions without discarding the underlying motion prior. In parallel, the object motion decoder predicts a target object motion aligned with the refined target character motion, ensuring that the object trajectory remains consistent with how the interaction is realized by the target character.

Learning Realistic Athletic Sprinting Without Demonstrations

4.0/5 相关 裁判分 5.0 Candidate 当日相对
William Wang, Nicholas Bianco, Guy Tevet, ... , Scott Delp, Kayvon Fatahalian
总结: 无演示肌肉驱动仿真生成真实竞技短跑。
方法: 高性能GPU仿真结合大批量强化学习训练。
证据: 生成近视觉真实的完整短跑与训练动作。
为什么适合我: 利于腿足高速运动与无演示强化学习。
推荐理由: 最接近物理角色控制与腿足运动先验:无示范强化学习在肌肉空间生成百米冲刺等高速全身运动。
原摘要

We present a muscle-driven simulation system for generating biomechanically accurate motion for high-speed athletic locomotion tasks that does not require motion demonstrations. Our approach integrates state-of-the-art biomechanical athlete models into a new, high-performance GPU simulator capable of running at 1000x real-time. High-throughput simulation enables large-batch reinforcement learning to train control policies that operate directly in the model's high-dimensional muscle excitation space, and are guided only by task-specific episode termination conditions and a reward that encourages maximizing speed while reducing forces needed to respect joint limits. These policies train within a few hours on a single GPU and generate "near visually realistic" motions for complete athletic activities such as a full 100-meter sprint or performing popular athletic locomotion drills like side-shuffling, backpedaling, and carioca. The generated sprinting motions also exhibit strong agreement with experimental data captured from sprinters.

Pengfei Zhang, Teng Sun, Xianchao Xiu
总结: 双潜空间强化学习优化生成式机器人策略。
方法: 预测噪声与表示潜变量并残差注入中间特征。
证据: 直接调制动作表示克服噪声引导局限。
为什么适合我: 融合扩散与强化学习提升全身策略。
推荐理由: 方法上接近扩散/生成式策略的强化学习引导采样,但对象是通用机器人动作潜空间,未落到人形全身或腿足运动。
原摘要

Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.

Giovanni B. Dessy, Claudio Semini, Victor Barasuol
总结: 被动机构影响四足负载携带不同步态。
方法: 仿真比较阻尼配置下爬行与小跑步态。
证据: 欠阻尼增加振荡并降低爬行ZMP裕度。
为什么适合我: 启发腿足接触丰富负载loco-manipulation。
推荐理由: 只分析四足负重被动臂的步态-阻抗与ZMP裕度,不涉及感知运动、跑酷或学习式全身控制。
原摘要

Passive mechanical interfaces offer a lightweight alternative to actuated manipulators for quadruped payload carrying, but their impedance directly couples the payload dynamics with the locomotion pattern. This paper analyzes how passive-arm stiffness-damping selection affects payload-carrying locomotion under different gait and payload conditions. We compare damped and underdamped passive-arm impedance configurations in simulation during flat-ground locomotion. For crawl gaits, where the support polygon remains well defined, the results show that underdamped impedance increases passive-joint oscillations and can reduce the ZMP margin with respect to the support polygon. Trot is retained as a dynamic excitation case for the passive arm, but it is not used for direct ZMP-margin stability comparison. The results are summarized in gait-payload-stiffness-damping maps, where ZMP-margin reduction is evaluated for crawl gaits and trot is retained only as a passive-arm excitation case.

Jiawen Wang, Kevin Yao, Khalid Jawed
总结: 障碍感知扩散策略泛化杂乱场景操作。
方法: 轻量障碍感知编码器提取结构化表示。
证据: 真实温室试验优于基线提升泛化。
为什么适合我: 扩散策略利于接触丰富障碍场景操作。
推荐理由: 障碍感知扩散策略用于末端轨迹与温室避障,属于臂式模仿操作,不是人形全身 loco-manipulation 或腿足运动。
原摘要

Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse trials per method (366 executions in total). ObstaDiff achieves 75.41% average task success and 8.20% average obstacle collision rate, outperforming representative imitation-learning baselines and improving generalization in cluttered agricultural scenes.

Multi-Modal Controlled Coherent Motion Generation

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Yifei Liu, Qiong Cao, Hongwei Yi, Huaiguang Jiang, Changxing Ding
总结: 多模态输入生成连贯逼真人形运动。
方法: 扩散模型独立生成各模态并组装部位。
证据: 无需对齐数据处理语音文本轨迹输入。
为什么适合我: 利于人体动作多模态生成与机器人重定向。
推荐理由: 语音、文本与轨迹驱动的虚拟人连贯动作生成,偏角色动画,不是可跟踪、可重定向、可上真机的全身控制。
原摘要

It is natural for humans to walk and talk simultaneously. This paper tackles the challenge of replicating such natural behaviors in 3D avatar motion generation driven by concurrent multimodal inputs, such as a text description of a man walking alongside speech audio. Existing methods, constrained by the scarcity of aligned multimodal data, typically combine motions from individual modalities sequentially or through weighted sums. However, they often result in mismatched or unrealistic movements. To overcome these limitations, we propose MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data. Our key innovation lies in decoupling the motion generation process. During each denoising step, the diffusion model independently generates motions for each modality from the input noise and assembles the body parts according to predefined spatial rules. The resulting combined motion is then diffused and serves as the input noise for the subsequent denoising step. This iterative approach enables each modality to refine its contribution within the context of the overall motion, progressively harmonizing movements across modalities. Consequently, the generated motions become increasingly natural and fluid with each iteration, achieving coherent and synchronized behaviors. We evaluate our approach using a purpose-built multimodal benchmark. Experimental results demonstrate that MOCO outperforms existing baselines, advancing the field of multimodal motion generation for 3D avatars.

Yaoyuan Yan, Zhiyou Heng, Haoxiang Jie, ... , Hongjie Yan, Wei Zhou
总结: 统一具身运行时实现四足闭环巡检。
方法: 分层组织运行时技能认知代理与交互。
证据: 共享上下文支持语音交互与知识复用。
为什么适合我: 支持腿足复杂场景感知运动闭环部署。
推荐理由: 四足社区巡检的智能体运行时与系统集成,不研究运动跟踪、跑酷或接触丰富全身控制。
原摘要

Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response within a traceable operational loop. Existing quadruped inspection systems commonly integrate these functions through task-specific interfaces, making contextual coordination, knowledge reuse, and controlled adaptation difficult. This paper presents \textit{Harness Robotic OS} (HROS), a unified embodied-agent runtime, and Argos, its realization for residential-community inspection. HROS organizes the system into robot runtime, embodied autonomy skills, cognitive agent runtime, and interaction and operations planes. A shared context connects physical state with agent reasoning; streaming ASR/TTS supports voice-based mission interaction; hierarchical working, episodic, and semantic memory preserves operational knowledge; and a safety-gated self-evolution loop converts execution traces into versioned candidate updates without permitting unconstrained online modification. The Argos prototype integrates a Vbot quadruped, Fast-LIO2 localization and mapping, Hobot-Stereo depth perception, PCT-Planner global planning, EGO-Planner local motion generation, and OpenClaw-orchestrated Qwen3-VL inspection analysis. Experiments in a residential property environment achieved 100\% waypoint reachability, outdoor localization error below 10~cm, local obstacle-response latency below 200~ms, representative hazard-detection rates of 85--95\%, and 99\% success in alarm delivery and structured-report generation. These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.

Hao Shi, Xi Li
总结: 拓扑必然性提供跨具身子目标控制。
方法: 同调从轨迹提取不可跳过阶段门控。
证据: 固定自由空间下独立于执行器存活。
为什么适合我: 启发人形腿足跨形态目标条件强化学习。
推荐理由: 跨本体目标条件强化学习的拓扑子目标理论,与人形感知运动、动作重定向无直接关系。
原摘要

Long-horizon goal-conditioned reinforcement learning delegates control to a high-level module that proposes subgoals, but existing subgoals are implicit byproducts of value functions or latent actions, tied to the executor that produced them. We study a different object: a route-conditioned order of unavoidable stages that every successful executor must traverse, recoverable from offline trajectories and belonging to none of them. Its defining properties are topological: an unskippable stage is a separating set that every admissible path must cross, and a loop in free space forces a route choice. We read the two by homology in dimensions 0 and 1 over a transport-weighted carrier built from successful trajectories, yielding an enumerable gate set with shell-level certificates; the certified gates are what we call topological necessities. Certified gates enter the decision loop as a recursive topological gate hierarchy. Under a fixed, isomorphic free space, the object survives executor replacement: gates frozen on PointMaze data transfer without retraining to Ant and Humanoid, attaining the highest Humanoid aggregate under a unified interface (96.1), with +36.0 over a map-privileged reference on the multi-route task (p=1.4e-5); the planner saturates PointMaze (100+/-0) and matches or exceeds the strongest baselines on AntMaze (giant +22.9) and Kitchen (+15.8/+12.6).

Kai Stewart, Yasunori Toshimitsu, Robert K. Katzschmann
总结: 实时雅可比估计快速学习手内笔写。
方法: 物理机器人估计手物系统任务雅可比。
证据: 短时初始化后在线适应无需模型演示。
为什么适合我: 接触丰富灵巧操作利于loco-manipulation。
推荐理由: 接触丰富的灵巧手在手书写与实时雅可比控制,不是全身人形控制或腿足 loco-manipulation。
原摘要

Dexterous in-hand manipulation of a grasped object with an anthropomorphic hand is an unsolved frontier for robot dexterity. The contact-richness and highly dynamic nature of object-hand interactions tend to require extensive modeling or data-collection efforts for learning-based approaches. Modern simulators used for reinforcement learning (RL) cannot fully replicate the required contact complexity, while collecting dexterous demonstrations for imitation learning (IL) remains an open problem. In this research, we present an embodied control approach based on real-time task Jacobian estimation of the combined hand and object system on the physical robot. Using only the CPU on a laptop, the proposed controller begins in-hand pen writing after approximately 18 s of initialization and continues to adapt online, without an analytic hand--object kinematic/contact model, simulation training, or precollected task demonstrations. We demonstrate that the same estimator/controller formulation works on three anthropomorphic robotic hand systems (one physical, two simulated) to show human-like, in-hand articulation of a grasped pen by an embodiment-independent formulation. Sub-millimeter in-plane precision (mean 0.6 mm across runs) is achieved across letters and shapes written in the air and on paper on a physical robot. To our knowledge, this is the first demonstration of an anthropomorphic hand writing arbitrary single-stroke trajectories with a grasped pen through purely in-hand motion, and it showcases an alternative to compute- and data-heavy approaches such as RL and IL for achieving dexterous manipulation through computationally simple and data-efficient algorithms.

A K M Nadimul Haque, Sheila Sutjipto, Marc G. Carmichael, Teresa Vidal-Calleja
总结: 安全引导强化学习适应动态环境技能。
方法: 高斯过程参数化顺序适应局部轨迹窗口。
证据: 产生连贯更新降低动作空间与信用分配难。
为什么适合我: 安全技能适应利于复杂地形接触强化学习。
推荐理由: 动态环境中的安全技能自适应强化学习,面向通用轨迹技能而非人形/腿足全身运动。
原摘要

Skill adaptation frameworks based on reinforcement learning often require restrictive assumptions to maintain stability, such as fixed observations or tightly controlled exploration schedules. In cluttered and dynamic environments, however, unrestricted exploration can lead to unsafe behaviour and unstable learning, particularly when task-relevant observations lie near obstacles or involve moving objects. In this work, we present Dist-GPRL, a distance-aware and safety-guided reinforcement learning framework for structured robot skill adaptation. Building upon Gaussian Process (GP)-based skill parameterisation, our framework sequentially adapts overlapping local windows of sparse trajectory via-points rather than modifying the complete skill at every policy step. Raw policy outputs are correlated through the GP covariance structure, producing temporally coherent trajectory updates while reducing the action-space and credit-assignment difficulties associated with global trajectory adaptation. Safety is incorporated through two complementary forms of guidance. A safe-subspace prior derived from the Hausdorff Approximation Planner (HAP) biases policy exploration toward feasible regions, while dynamically updated distance field clearance and gradient rewards provide local obstacle awareness. A trajectory-kinematics similarity regulariser further preserves the demonstrated velocity and acceleration characteristics during adaptation. We evaluate the framework on two dynamic object-manipulation tasks in simulation and transfer the learned policy to real-world robot execution. Experimental results demonstrate higher task success, lower collision frequency, and more stable learning than the baselines, while preserving the kinematic characteristics of the demonstrated skill.