Papers for 2026-10-07

10 papers

Humanoid Loco-Manipulation With Discrete VLA Model

5.0/5 很相关 裁判分 10.0 Top pick 当日相对
Wenxin Shao, Siqi Chai, Kun Li, ... , Wei Xu, Qiang Liu
总结: 首个离散VLA模型实现人形loco-manipulation。
方法: 统一分词器分解四部位动作并扩展语言模型词汇。
证据: 跨遥操作、人类视频与仿真数据训练不同本体。
为什么适合我: 直接支持接触丰富场景下全身loco-manipulation。
推荐理由: 直接做人形全身 loco-manipulation 的离散 VLA,并融合遥操作、第一视角人体视频与仿真数据,紧贴全身 loco-manipulation、遥操作数据采集与模仿学习。
原摘要

Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.

Xiaoyu Yang, Sen Han, Da Li, Nan Wu
总结: 单目视频几何保持人体到机器人上身重定向。
方法: 统一体手重建与形态无关几何转移加多阶段IK。
证据: 处理尺度运动学差异并保留精细远端运动。
为什么适合我: 从视频提取可跟踪可重定向上身全身控制。
推荐理由: 从单目视频做人到机器人的上身与手部几何保持重定向,直接对应人体动作重定向、模仿与跨本体迁移。
原摘要

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer. Frame-wise body estimates, video-level observations, and detailed hand evidence jointly constrain a single differentiable Momentum Human Rig (MHR) state, while transient hand artifacts are repaired in parameter space. The reconstructed motion is represented by arm-segment directions, elbow configuration, relative palm orientation, and bilateral wrist relations, and is realized on the target robot through multi-stage inverse kinematics and robot-specific hand adaptation. Within the broader system, Across-VAM provides video generation, whereas Across-WAM performs human-to-robot motion mapping. The method is evaluated on 16 monocular videos comprising 1,769 source frames, including 10 signing and six reach-to-grasp sequences. Unified reconstruction reduces mean hand reprojection error from 22.36 to 7.21 pixels relative to SAM 3D Body. All 16 retargeted trajectories completed kinematic simulation playback, and representative signing and reach-to-grasp motions were further demonstrated on a physical robot. The results demonstrate a unified pipeline from monocular human video to coordinated upper-body motion on a dual-arm dexterous robot.

Retargeting Motions to Diverse Skeletons via Learnable Flattening

4.0/5 相关 裁判分 9.0 Top pick 当日相对
Kia-Jüng Yang, Fabian H. Sinz, Paweł A. Pierzchlewicz
总结: 可学习展平实现多样骨架跨结构运动重定向。
方法: Transformer自编码器学习拓扑平移不变潜空间。
证据: 无监督处理未见拓扑,优于几何与Transformer方法。
为什么适合我: 助力人体动作重定向到异构人形机器人骨架。
推荐理由: 跨骨架拓扑的动作重定向,对应人体动作重定向与物理角色控制;偏图形学零样本骨架迁移,而非真机全身跟踪。
原摘要

Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by $43-47\%$ over current benchmarks. A user study ($n = 37$), including expert animators, further ranks our approach highest in motion alignment and physical plausibility ($p < 0.05$). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.

ReFM: Semantic-Aware Refinement Flow Model for Motion Retargeting

4.0/5 相关 裁判分 9.0 Top pick 当日相对
Jingxiang Qu, Lucie Taglienti, Evan Atherton
总结: 语义感知精炼流模型解决运动重定向问题。
方法: 无配对数据学语义,流模型精炼避免伪影纠缠。
证据: 平衡语义保真与目标物理时间约束。
为什么适合我: 提升人体动作到机器人重定向的语义物理保真。
推荐理由: 用流模型做跨骨架语义动作重定向并兼顾物理合理性,贴近人体动作重定向与生成式运动先验,但是角色动画而非机器人控制。
原摘要

Motion retargeting transfers motion across characters with different skeletal structures while preserving semantic intent and physical plausibility. Despite recent progress, two fundamental questions remain: (i) how can reliable source-motion semantics be learned without high-quality paired retargeting data, and (ii) how should retargeting be formulated when no reliable paired motion can serve as a definitive regression objective? Existing methods commonly preserve semantics by constraining predictions toward copied motions. However, such initializations entangle useful articulation cues with artifacts caused by mismatched skeletal proportions and body geometry. Moreover, directly regressing a final motion in one forward pass is restrictive because retargeting is inherently underdetermined, and the desired solution must balance semantic fidelity with target-specific physical and temporal constraints rather than match a unique paired target. Motivated by these limitations, we propose ReFM, a source-mesh-agnostic, energy-guided model that reformulates motion retargeting as progressive refinement. First, an SO(3) canonicalizer removes redundant global-orientation variations. Second, a cross-character semantic encoder, pretrained through contrastive learning, provides a character-invariant representation for both optimization guidance and semantic evaluation. ReFM then progressively refines an initialized target motion through a learned flow guided by semantic consistency, physical plausibility, temporal coherence, and minimal motion modification. The framework is compatible with different initialization strategies, including both direct motion copying and Autodesk HumanIK, an industry-standard full-body inverse-kinematics retargeting system.

Qichen Zheng, Siyuan Yang, Chong Wang, ... , Alex Kot, Kwok-Yan Lam
总结: 拓扑感知文本驱动编辑异构人形骨架运动。
方法: 流匹配变换器加拓扑约束传播与编辑对齐。
证据: 处理不同关节数层次并保留源运动内容。
为什么适合我: 支持文本编辑人体动作以适配多样机器人。
推荐理由: 面向异构人形骨架的文本驱动动作编辑,对应动作重定向与角色运动生成;仍是动画骨架编辑,不是全身控制或上真机。
原摘要

Text-driven motion editing modifies an existing motion sequence according to a text instruction while preserving the content of the source motion. Existing methods are typically built for a single, fixed skeletal topology, which limits their use in animation pipelines where characters differ in joint count and skeletal hierarchy. We present Topology-Aware Motion Editor (TAME), a flow-matching transformer that edits motions on humanoid skeletons of varying topology. TAME represents motion as per-joint, per-frame tokens and models interactions among joints, across frames, and with the text instruction through skeletal, temporal, and text cross-attention layers. To make the skeletal attention follow each character's hierarchy, TAME replaces full joint attention with Topology-Constrained Skeletal Propagation (TCSP), which restricts attention to one-hop kinematic neighbors in the skeleton's adjacency matrix. We further introduce Edit-Focused Representation Alignment (EFRA), a self-distilled representation alignment strategy that aligns student features with cleaner EMA-teacher features exclusively on edit-relevant joint-time tokens, making edits faithful to the instruction. To make this setting trainable and comparable, we construct TopoMotionFix, a multi-topology extension of MotionFix with seen- and unseen-topology evaluation protocols. TAME outperforms previous methods in edit alignment and source preservation on MotionFix and reliably edits motions on unseen skeletons in TopoMotionFix.

DexForge: High-Fidelity Physics-Informed Dexterous Retargeting

4.0/5 相关 裁判分 8.0 Strong 当日相对
Meizhong Wang, Kun Cao, Ruiqi Ni, Lihua Xie, Yiguang Hong
总结: 物理知情框架将人类视频转为高保真机器人轨迹。
方法: 可微仿真结合接触运动学与力动力学重定向。
证据: 在DexYCB和HOT3D演示跨七种灵巧手验证。
为什么适合我: 实现接触丰富灵巧操作的高保真重定向。
推荐理由: 从人类视频做物理一致的灵巧操作重定向,贴近跨本体运动重定向,但侧重手-物接触而非腿式全身运动。
原摘要

Human demonstrations offer rich examples of precise dexterous manipulation and a promising source of robot training data. However, high-fidelity reproduction of demonstrated motions and hand-object interactions across robot embodiments remains challenging under physical constraints. We present DexForge, a differentiable physics-grounded framework for converting human video demonstrations into high-fidelity robot trajectories. We reconstruct spherical-Gaussian object models and hand-object motion from visual observations, then build a differentiable simulator combining efficient Gaussian collision detection with existing differentiable dynamics. Based on this simulator, DexForge combines contact-aware kinematic retargeting with force-aware dynamics retargeting: robot-adapted stable contacts guide kinematic reference construction and subsequent gradient-based control refinement for precise physical motion reproduction. Experiments on 130 DexYCB and HOT3D demonstrations across seven dexterous hands show success-rate gains of approximately 35-53 percentage points over the baseline, with object position and orientation tracking errors on successful trajectories reduced by approximately 34-67% and 71-78%, respectively. Further experiments demonstrate open-loop transfer to MuJoCo and real-robot execution. Our project page is available at https://wmz1226.github.io/DexForge/

VAMPS: Visual and Motor Policies from Sampling-Based Planning

4.0/5 相关 裁判分 7.0 Strong 当日相对
Mohamed Yassine Kabouri, Pietro Noah Crestaz, Quang-Nam Nguyen, ... , Ludovic Righetti, Nicolas Mansard
总结: 采样规划框架训练视觉运动策略无需演示。
方法: MPPI迭代监督结合终端价值与IQL批评家。
证据: 迭代优于冻结数据并迁移到Unitree Go2。
为什么适合我: 促进腿足机器人安全无演示强化学习策略。
推荐理由: 用采样规划监督策略并迁移到Unitree Go2,贴近四足运动学习与仿真到真实,但框架本身更通用。
原摘要

Learning robot policies directly on physical systems remains difficult because data collection is costly and policy exploration can be unsafe. We introduce Visual and Motor Policies from Sampling-Based Planning (VAMPS), a framework that uses Model Predictive Path Integral (MPPI) control to train reusable policies without human demonstrations. VAMPS supports two training modes. For one-step proprioceptive policies, it operates iteratively in simulation: the policy warm-starts MPPI, and the refined trajectories provide new supervision as the policy changes. A learned terminal value improves short-horizon planning, while an Implicit Q-Learning (IQL) critic guides the policy update. Iterative refinement outperforms training once on frozen MPPI data, and we transfer the learned locomotion policy to a Unitree Go2. For visuomotor policies, VAMPS operates directly from real-robot data. MPPI uses task-specific state estimates to plan and execute trajectories while recording RGB and sensor observations on a Flexiv Rizon 10S. Action Chunking with Transformers predicts action chunks, reducing the effective prediction horizon, and is trained offline on this fixed dataset. We demonstrate visuomotor pick-and-place and force-aware whiteboard erasing. In the latter task, the policy additionally observes the measured $6$-D wrench and desired normal force. These results show that VAMPS can learn policies either in simulation followed by hardware transfer or directly from autonomously collected real-robot data.

Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S. K., ... , Abhay Dwivedi, Amit Shukla
总结: 分层强化学习实现欠驱动双足无碰撞运动。
方法: 高层速度命令低层跟踪联合SAC训练。
证据: 与A*、RRT*、APF基线在随机试验中对比。
为什么适合我: 适用于复杂地形接触场景下避障运动智能。
推荐理由: 欠驱动双足的分层强化学习避障行走与腿式运动有关,但平台简化且重点是无碰撞速度跟踪,不是接触丰富的全身敏捷控制。
原摘要

A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, \omega_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.

ROOT: Discovering Rewards for User-Specified Embodied Behaviors

3.0/5 一般 裁判分 7.0 Strong 当日相对
Eren Sadikoglu, Aditya Taparia, Xinyuan Liu, Ransalu Senanayake
总结: 观察树框架发现奖励匹配用户指定行为。
方法: 视频语言模型诊断失败引导奖励程序搜索。
证据: 捕捉自然步态姿势等视觉易识别行为属性。
为什么适合我: 优化RL奖励生成自然可部署全身运动风格。
推荐理由: 以自然步态、姿态和风格为例发现奖励,与风格化运动相邻,但本质是通用奖励搜索而非腿式控制器。
原摘要

Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over a persistent experiment tree that stores reward programs, trained policies, and rollout observations, together with behavioral insights distilled by a video-language model, to diagnose behavioral failures and guide subsequent reward refinements. We evaluate ROOT on seven tasks across four embodiments: simulated Hopper, HalfCheetah, Ant, Unitree Go2, and as well as the real-world Unitree Go2. ROOT produces behaviors that better align with user intent than those generated by existing LLM-based reward-generation methods, achieving up to 86.8% locomotion-completeness accuracy and improving Vid-LLM behavioral alignment from 3.56/5 to 4.14/5, a 16.5% improvement over baselines. Human evaluations further support these results, with ROOT preferred in 51-63% of pairwise comparisons.

Duc Cuong Vu, Van Tung Nguyen, Duc Hai Nguyen, ... , Vu Trung Tran, Minh Nhat Vu
总结: 共形种子混合逆解求解偏移冗余人形臂。
方法: 解析候选加共形预测排序后数值精炼。
证据: 实时高精度克服初始化敏感与近似限制。
为什么适合我: 支持人形机器人实时全身冗余臂控制。
推荐理由: 人形级冗余机械臂的实时混合逆运动学,可服务上肢控制,但只是运动学求解器,不是全身运动智能或学习控制。
原摘要

This paper presents a conformal-seeded hybrid strategy for solving inverse kinematics of offset, redundant 7-DoF robot arms of the humanoid class. Analytical inverse kinematics (AIK) provides closed-form solutions with very low computational cost. However, for offset kinematic structures, the exact closed-form solution is generally unavailable, and practical AIK must rely on an approximate or simplified kinematic model. In contrast, numerical inverse kinematics (NIK) can achieve high-precision solutions on the full kinematic model. However, its convergence is highly sensitive to initialization. To overcome these limitations, we propose a two-stage hybrid inverse kinematics framework with conformal-calibrated seed selection. First, an approximate analytical model efficiently enumerates a finite set of candidate joint solutions. Second, we rank these candidates using a lightweight learned predictor of post-refinement difficulty, wrapped by split-conformal prediction into a calibrated upper bound that serves as the selection score. The best-ranked seed is then refined using a Levenberg-Marquardt solver on the full kinematic model. The proposed method combines fast candidate generation, learned seed ranking with a calibrated difficulty bound, and accurate numerical refinement, achieving real-time performance of less than 40us and a success rate of 100% in our evaluation on reachable targets. We validate the approach through large-scale stochastic simulation across the workspace and experimental demonstrations with motion planning on a humanoid robot arm. Demonstration videos are available at https://youtu.be/aeiBmw1XRbw.