Papers for 2026-09-15

10 papers
Yilin Zou, Chenghua Liu, Chenglong Wu, Fanghua Jiang
总结: 流匹配运动先验用在线OT奖励改进模仿学习。
方法: 熵OT耦合路径,流匹配训势函数得标量奖励。
证据: 受控实验显示奖励模型泛化显著更好。
为什么适合我: 改进AMP先验,利于人形全身运动模仿控制。
推荐理由: 直接改进对抗运动先验:用在线最优传输与flow matching给模仿学习提供运动先验奖励,并明确针对步态相位,最贴近物理角色控制与运动先验。
原摘要

Learning a motion prior requires a reward that guides a policy from its current behavior toward demonstrated motion. Adversarial Motion Priors (AMP) provide such a reward with a discriminator. However, adversarial objectives can become uninformative when policy and expert supports are far apart. A naive use of optimal transport (OT) averages matched expert successors into a barycentric target. Averaging across gait phases can weaken the target's joint motion. We introduce Flow-Matched Motion Priors (FMP), an online scalar reward learned from paths connecting current rollout histories to an expert motion bank. Entropic OT supplies the coupling. Before each policy update, we train a neural potential with flow matching (FM) along the rollout-to-expert paths, endpoint-gradient supervision, and relative-value calibration. The actor receives only physical observations and the reward remains a scalar, as in AMP. Controlled reward-model experiments show substantially better generalization beyond the fitting rollout than value-only or endpoint-only fitting. On Unitree G1, matched 50-million-transition experiments compare FMP with AMP, a barycentric OT reward, and nested ablations under demonstration and fixed-pose initialization. FMP produces stable forward walking at 0.727 m/s from demonstration resets and 0.338 m/s from a fixed default pose. In the fixed-pose condition, it incurs 129 falls versus 243 for the endpoint-only control. Against a static score-gradient teacher, dynamic FM reduces score-increment error at interpolation fractions 0.25 and 0.50 while using 29% less offline fitting time.

EMoG: Emotion-Modulated Gait Generation for Expressive Humanoid Locomotion

4.0/5 相关 裁判分 9.0 Top pick 当日相对
Yi Lu, Tianhao Jiang, Honglong Tian, ... , Qiu Shen, Xun Cao
总结: 情感调制步态实现表达性人形运动。
方法: 风格码条件MLP生成步态,统一RL策略跟踪。
证据: 实验实现连续步态风格调制与交互控制。
为什么适合我: 增强人形表达运动,结合RL支持全身控制。
推荐理由: 面向人形步态生成与统一强化学习跟踪,并用真人步态数据,贴近人形全身控制与运动跟踪;侧重情绪表现而非复杂地形或跑酷。
原摘要

Existing humanoid locomotion systems primarily focus on stability and task execution, while integrating expressiveness with explicit locomotion control remains challenging. We propose EMoG, an emotion-modulated gait generation framework for expressive humanoid locomotion. EMoG introduces an emotional-style code with continuously adjustable intensity. Conditioned on this code and physical commands, a lightweight MLP generates expressive, command-consistent periodic gait trajectories in real time, which are tracked by a unified reinforcement learning policy for physical execution. To support training, we collect a large-scale emotion-annotated gait dataset from professional performers and develop an automated pipeline to extract physically consistent periodic gait cycles. EMoG also integrates an LLM-based parser that converts free-form language into emotional style and motion parameters for interactive control. Experiments demonstrate that our system achieves continuous gait-style modulation with perceptible expressive cues while maintaining command tracking. EMoG provides a practical approach to parameterized emotional-style walking for human-robot interaction.

Liang Zhou, Jiaming Su, Yancong Wei, Kangkang Dong, Houde Liu
总结: 退化视觉下四足操作器鲁棒移动抓取。
方法: 教师学生框架结合抓取推理与时序估计。
证据: 基准评估退化视觉下成功效率与平滑。
为什么适合我: 复杂地形接触感知,利于腿足loco-manipulation。
推荐理由: 四足操作臂在复杂地形上的移动抓取与全身控制,贴近腿足感知运动和loco-manipulation;重点是退化视觉下的抓取而非运动智能本身。
原摘要

Quadruped manipulators enable mobile grasping in complex environments, yet their whole-body control policies remain vulnerable to unreliable onboard visual perception. Existing methods are typically developed under relatively reliable observations and have not systematically examined how occlusion, segmentation-mask dropout, depth noise, and target-localization jitter affect grasp reasoning and target tracking. To address this gap, we introduce DeViGrasp-Bench, a benchmark for mobile grasping under degraded vision that incorporates controlled visual degradations, seen and unseen objects, multiple difficulty levels, and complex terrains, and evaluates task success, execution efficiency, and action smoothness. We further propose DeViGrasp-Net, a teacher--student framework that combines state-conditioned grasp reasoning with reliability-aware temporal target estimation. The privileged teacher attends to offline grasp candidates conditioned on object, robot, end-effector, and task states, while the deployable student fuses dual-view segmented-depth observations with current, memory, and recovery target hypotheses through Target Hold Memory and Temporal Memory Attention. DeViGrasp-Net outperforms VBC across degradation levels, unseen objects, and complex terrains, and surpasses an adapted DQ-Net across all evaluated degradation levels. Under the Difficult setting, it achieves a success rate of 62.3\%, improving upon VBC and DQ-Net by 16.1 and 4.3 percentage points, respectively; under the Hard setting, its margin over DQ-Net increases to 10.5 percentage points. Ablation studies confirm the complementary benefits of grasp-aware supervision and reliability-aware temporal memory.

Jiahao Liu, Kento Kawaharazuka, Tasuku Makabe, Kei Okada
总结: 连续流形阻抗重定向用于接触模仿学习。
方法: 扩展为连续变阻抗,自动编译求解结构。
证据: 真实接触任务试验改进保留与力指标。
为什么适合我: 接触丰富模仿,支持全身控制可重定向。
推荐理由: 接触丰富模仿与阻抗重定向方法相邻,但未体现人形或腿足全身运动、运动跟踪或上真机全身控制。
原摘要

CMDIR extends Manifold-Decomposed Impedance Retargeting (MDIR) to transform fixed-impedance demonstrations into continuous variable-impedance controllers, which can also serve as structured supervision for imitation learning. Continuous Task-Manifold Impedance Representation (TMIR) pairs an evolving task frame with controller instructions. Demo-relative Compromise dynamics retain moving-basis transport and control/physical metric mismatch, yielding displacement, reaction-impulse, and perturbation-sensitivity criteria. Quality-to-Fast automatically compiles a solver structure from development paths within a predefined finite space, re-instantiates that structure for each demonstration, and certifies the resulting candidate by multi-resolution evaluation. Across 225 retargeted-controller trials in three real contact tasks, full CMDIR improves mean task-proxy retention and reduces mean pose deviation, force fluctuation, and peak force relative to discrete MDIR. FastMPO achieves a $5.8$--$9.4\times$ speedup over C-MPO with comparable closed-loop outcomes. Downstream experiments demonstrate learnability of the complete TMIR supervision interface; lower force fluctuation and peak force are observed among successful executions, while completion reliability remains uneven across tasks and environments.

Emek Barış Küçüktabak, Karankumar Patel, Jinda Cui, ... , Kazuhiro Sasabuchi, Jun Takamatsu
总结: 原语引导采样MPC实现多指灵巧操作。
方法: 低维原语偏置采样并优化关节残差。
证据: Allegro手消融证明原语与残差必要。
为什么适合我: 全身MPC接触操作,利于灵巧manipulation。
推荐理由: 采样式模型预测控制与接触丰富操作相邻,但对象是多指灵巧手,不是腿足或人形全身控制。
原摘要

We present a primitive-informed sampling-based model predictive control (MPC) framework for multi-fingered dexterous manipulation. Sampling-based MPC evaluates candidate control trajectories through forward simulation without requiring gradients through complex contact dynamics. However, directly sampling these trajectories in the high-dimensional joint space of a dexterous hand is inefficient and makes performance strongly dependent on the sampling distribution. Our framework biases sampling using low-dimensional manipulation primitives that encode coordinated finger motions, while simultaneously optimizing joint-level residuals to adapt these motions to the current hand-object configuration. Task-related rollout constraints reject infeasible trajectories during forward simulation, improving the effective use of the sampling budget. We evaluate the approach on a 16 DoF Allegro hand using a synchronized MuJoCo digital twin. Ablations show that both the primitive and residual are necessary for reliable continuous in-hand rotation, that increasing the sampling budget alone does not recover this coordination, and that rollout constraints substantially improve success rate. A primitive extracted for one object size transfers to other sizes and remains effective under model mismatch. The framework further supports grasping, object reorientation, and coordinated arm-hand manipulation, using primitives extracted from both a simulation-trained policy and human hand-motion data.

Guocun Wang, Kenkun Liu, Guorui Song, ... , Xiaoguang Han, Haoqian Wang
总结: 开放世界统一运动语言理解与生成。
方法: 统一token空间加运动一致思维链训练。
证据: 百万数据促进模态平等与长序列生成。
为什么适合我: 人体动作生成理解,支持可跟踪重定向。
推荐理由: 开放世界人体动作—语言生成与理解,只弱相关于动作先验,没有机器人跟踪、重定向或全身控制。
原摘要

Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-modal interaction. Moreover, the next-token prediction paradigm is not naturally suited to long motion sequences, where autoregressive generation may accumulate prediction errors. To address these challenges, we propose Open-UniMo, a unified Large Motion-Language Model (LMLM) trained on million-scale open-world motion-language data. Open-UniMo promotes modality parity by extending Qwen's vocabulary of about 150K text tokens with 64K motion tokens, enabling motion and language to share a unified token space. We further introduce motion-consistent Chain-of-Thought reasoning as an intermediate representation to bridge language semantics and motion dynamics. Open-UniMo is trained with a two-stage pipeline, where supervised fine-tuning establishes CoT-guided bidirectional motion-language mapping and Group Relative Policy Optimization (GRPO) improves semantic alignment while mitigating cumulative errors in autoregressive motion-token generation. To support comprehensive evaluation, we propose Open-MoBench, a VLM-guided benchmark for assessing text-to-motion (T2M) generation, motion-to-text (M2T) understanding, and bidirectional consistency. Extensive experiments show that Open-UniMo achieves state-of-the-art performance on both conventional metrics and Open-MoBench. Furthermore, ablation studies reveal that M2T understanding is not primarily limited by motion-token vocabulary size; instead, coupling M2T with the learnable T2M generation path yields stronger cross-modal representations, demonstrating that generation can facilitate understanding in AR-based motion-language modeling.

MoVT: Video-Augmented Motion Tokenizer for Text-to-Motion Generation

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Beibei Jing, Tianle Guo, Youjia Zhang, ... , Tao Guan, Wei Yang
总结: 视频增强tokenizer提升文本到运动生成。
方法: 跨模态投影丰富码本,掩码transformer预测。
证据: 利用视频丰富复杂真实运动模式表示。
为什么适合我: 人体动作生成先验,利于模仿与重定向。
推荐理由: 视频增强的文本生成三维人体动作,仅弱相关于动作数据,不是可跟踪的机器人全身控制。
原摘要

Text-driven 3D human motion generation models face significant challenges in responding to diverse and unconstrained textual prompts, primarily due to the limited availability of 3D motion training data. To address this, we introduce MoVT, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation. At the core of our approach is the cross-modal augmented motion tokenizer, which projects discrete 3D motion tokens into the 2D domain. This projection allows us to enrich the motion codebook with complex, real-world motion patterns derived from videos. The enriched discrete tokens are then mapped back to the 3D domain, resulting in aligned 3D and 2D codebooks with an enhanced capacity to represent intricate motions. These enhanced codebooks are integrated into a generative masked transformer, which predicts masked motion token indices in a modality-agnostic manner. This enables the use of text-index pairs, generated from the 2D codebook and annotated motion videos, to further enhance the generator. Extensive empirical evaluations show that MoVT performs favorably against prior state-of-the-art methods across multiple key metrics.

Touch2Trace: Tactile-Driven Imitation Learning for Dexterous Cable Tracing

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Matteo Grimaldi, David Klee, Ziling Chen, ... , Tao Yu, Saleh Nabi
总结: 触觉驱动模仿实现灵巧电缆追踪。
方法: 自监督触觉编码加transformer行为克隆。
证据: 触觉反馈显著优于仅本体感觉追踪。
为什么适合我: 接触丰富灵巧操作,触觉感知运动控制。
推荐理由: 触觉驱动的灵巧手线缆追踪模仿学习,属于手部接触操作,而非腿足全身或loco-manipulation。
原摘要

Dexterous manipulation of deformable objects demands continuous fingertip-level regulation of pressure, friction, and incipient slip. We study one of the most challenging cases: dexterous cable tracing, feeding a cable through the hand with repeated pinch-and-curl motions of the thumb and index finger. We introduce Touch2Trace, a tactile-driven imitation-learning system for this task, and provide, to our knowledge, the first systematic real-world characterization of how encoder pretraining, control rate, temporal context, and spatial resolution each shape policy performance. The winning learning recipe combines a tactile encoder pretrained for a custom 32 x 32 piezoresistive sensor (TacV5) via self-supervised learning with a lightweight transformer policy trained on teleoperated demonstrations via behavior cloning, deployed at 60 Hz on a Tesollo DG-5F hand. Tactile feedback without vision or explicit cable-state estimation significantly improves tracing performance versus a proprioception-only baseline: from 0.2 cm to 20.1 cm mean distance and 0% to 93% success rate, with zero-shot transfer to unseen cables and routing conditions. The results quantify the influence of key parameters in tactile-driven systems for reliable dexterous deformable object manipulation.

OJOx: Specification-Conditioned Demonstrations for Embodied AI in Construction

2.0/5 偏低 裁判分 4.0 Candidate 当日相对
Mohamed Dawod
总结: 规范条件演示用于建筑具身智能。
方法: 同步感知状态设计意图与行为捕获。
证据: 为施工提供规范条件演示数据接口。
为什么适合我: 人体演示可重定向,支持复杂场景模仿。
推荐理由: 建筑场景的规格条件化人体示范采集,只弱相关于数据采集,没有机器人重定向或运动控制。
原摘要

Large-scale egocentric and whole-body human demonstrations are becoming a primary source of data for embodied intelligence. They record what people perceive and do, but rarely the external specification that gave an action its purpose. In construction that omission is consequential: skilled work is directed at project-specific configurations defined in a design model - configurations not yet present in the environment being observed. A mason's transferable competence is not the geometry of one wall but the ability to realise a new geometry from a specification. We introduce the specification-conditioned demonstration: a synchronised record of the physical state a demonstrator perceives, the intended state supplied to them by an external design, and the behaviour connecting the two. We present OJOx, a capture interface that realises this for construction - delivering design geometry to a headset, anchoring it in the physical workspace, rendering it into a demonstrator's stereo passthrough view, and recording that view synchronously with whole-body and hand motion. We report one fully instrumented session - a 33-component wall laid against a specification that changes while the work proceeds - and check the record against the physical scene through an external camera registered independently of the capture. Recorded sessions remain compatible with existing humanoid retargeting infrastructure and replay onto a Unitree G1 in simulation. The result is a data interface for testing whether embodied policies can learn not merely to imitate demonstrated actions, but to act toward specifications absent from their training experience.

Learning In-Hand Object Reaching to General 6D Poses

2.0/5 偏低 裁判分 4.0 Candidate 当日相对
Junxiao Lin, Tianyue Wu, Jie Yin, ... , Kaifeng Zhang, Weiming Zhi
总结: 学习手内物体到达通用6D姿态。
方法: 多样抓取初始化加自适应6D目标课程。
证据: 仿真提升持有抓取成功与掉落后恢复。
为什么适合我: 灵巧手内操作,接触丰富全身运动智能。
推荐理由: 手内6D物体位姿到达的仿真到真机强化学习,不是人形或腿足全身运动。
原摘要

In-hand manipulation allows multi-fingered dexterous hands to reconfigure grasped objects without releasing and regrasping them. This improves manipulation efficiency by reducing repeated grasp acquisition and large arm motions. However, most learning-based methods focus on reorientation, continuous rotation, or translation, whereas many tasks require joint control of object position and orientation. We formulate this capability as in-hand 6D object pose reaching: starting from an existing grasp, coordinated finger motions move the object to a palm-relative target pose. We present POISE (Palm-relative Object reaching In SE(3)), a sim-to-real reinforcement learning framework for this task. POISE combines diverse stable-grasp initialization, goal- and geometry-conditioned control, an adaptive 6D goal curriculum, and a compact reward scheme for pose reaching and grasp preservation. In simulation, diverse initialization raises held-out-grasp success from 40.1% to 51.5% and post-drop recovery from 33.8% to 72.9%; the curriculum raises full-range success from 6.2% to 59.5%. On hardware, the grasp-maintenance reward improves three-target sequence success from 20% to 80%. In real-world experiments, POISE reaches successive 6D targets without manual reset across multiple object geometries and wrist orientations, and recovers from external disturbances. To support further research in dexterous manipulation, we will release our code at https://junxiaolin.github.io/poise-website/.