Papers for 2026-09-30

10 papers
Naichuan Sun, Haotian Shen, Yizhang Zhang, ... , Yaochu Jin, Peidong Liu
总结: 提出从人类演示学习灵巧人形移动操作的统一框架,同时迁移动作本身及其身体-手腕-手指-物体交互耦合结构。
方法: 两阶段重定向先分别初始化身体与手部动作,再在保留下肢支撑下对上肢交互链做耦合细化;跟踪策略用结构化解剖token与有向掩码注意力的Transformer。
证据: 摘要显示交互一致的参考可被解剖感知策略稳定跟踪,在灵巧移动操作上优于基线,但具体指标在摘录中被截断。
为什么适合我: 与我的运动重定向加全身RL路线直接相关,交互一致性参考生成与区域化注意力跟踪可借鉴到接触丰富移动操作的全身控制器。
原摘要

Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher-student distillation, or subsequent residual refinement. DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world. See our project page (https://dexweave.github.io) for videos.

Victor Paredes, Ayonga Hereid
总结: 将ALIP步态模板的离散步到步安全证书DECBF嵌入双足强化学习,兼作训练塑形信号与运行时摆动脚落点最小修正滤波器。
方法: 推导ALIP步行的离散指数控制屏障函数安全约束,训练时作为奖励塑形项,部署时以最小改动调节摆动脚落点满足模板级安全。
证据: 在MuJoCo中结合全身控制器于Digit人形实证评估,外部扰动试验中安全违例事件较无约束基线减少。
为什么适合我: 提供了把解析模板安全证书与RL全身控制融合的范式,可提升我双足敏捷行走工作在真实部署时的安全性与可解释性。
原摘要

Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.

Alberto Sanchez-Delgado, João Carlos Virgolino Soares, Victor Barasuol, Claudio Semini
总结: 面向类月环境的四足自主勘探框架,融合外感知几何可通行性映射与本体感知接触线索进行风险感知的自主探索。
方法: 机载RGB-D构建机器人中心高程图并导出导航代价,本体测量补充交互感知地形线索,增量配准为全局多层地图供探索模块选择目标。
证据: 在类月场景中实现自主目标选择与碰撞感知导航至未探索区域,系统展示了外-本体融合映射对不规则地形勘探的增益。
为什么适合我: 呼应我关注的地形利用与感知运动,其外-本体多层地图可作为四足/人形在非结构化地形上全身运动控制的上游感知模块。
原摘要

Autonomous planetary exploration requires robots to navigate unknown, uneven terrain while assessing risk, traversability, and energetic cost. Quadruped scouts are well suited for this task because they can traverse irregular surfaces and gather mobility-relevant information during locomotion. This paper presents a terrain-aware exploration framework that combines exteroceptive and proprioceptive mapping for a quadruped robot in lunar-like environments. An onboard RGB-D camera builds robot-centered elevation maps, estimates geometric traversability, and derives navigation costs for autonomous planning. In parallel, proprioceptive measurements provide interaction-aware terrain cues that complement geometry-based assessment. Local maps are incrementally registered into a global multi-layer representation, which is used by an exploration module to select targets in unexplored regions of interest. The targets are reached by an autonomous navigation system that guides collision-aware motion using the available map and cost layers. Simulation results on NVIDIA Isaac Sim show autonomous exploration, map expansion, and spatial association between terrain geometry and robot-terrain interaction. Subsequent navigation using this information exhibits lower average Cost of Transport (CoT) than initial exploration.

Jihwan Shin, Adrià López Escoriza, Junzhe He, Matthias Heyrman, Marco Hutter
总结: 提出以接触为中心的人-物交互重定向方法,把HOI迁移到人形机器人,从而大规模生成兼容的全身交互参考。
方法: 滑窗轨迹优化把每个标注接触作为物体坐标系下的目标,在机器人运动学极限下平衡身体跟踪、足部支撑与平滑性。
证据: 可跨物体尺寸增强单条演示、吸收单目视频重建的接触并扩展到多机器人协作,公开了代码与重定向运动数据集。
为什么适合我: 直接服务我的运动重定向需求,以接触为锚的优化目标可为我批量生成接触丰富移动操作的大规模全身训练参考。
原摘要

Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method that transfers HOI onto a humanoid robot for large-scale motion-data generation. Its windowed trajectory optimization uses every labeled contact as a target in the object frame, balancing body tracking, foot support and smoothness under the robot's kinematic limits. The method can augment a single demonstration across object sizes, absorb contacts reconstructed from monocular video, and extend to several robots manipulating one object. We publicly release the code and the retargeted motion dataset.

Zihao Wang, Shutong Liu, Siqi Zheng, ... , Yanchao Yang, Mengdi Xu
总结: 将全身分布式触觉引入VLA策略,构建触觉锚定的多模态上下文,解决接触区域被遮挡时的人形移动操作控制。
方法: 增设触觉通路,其隐状态不仅用于动作生成,还预测未来触觉、本体与视觉表征,以预测目标塑造结构化的物理世界理解。
证据: 在五个真实机器人任务上评估,涵盖触觉触发行走、持续物理交互与人机交互场景,验证全身体感适应的有效性。
为什么适合我: 为我的接触丰富全身控制补充触觉模态思路,预测式触觉潜表征可增强策略对不可见接触的响应能力与真机迁移。
原摘要

Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions, distributed tactile sensing preserves spatially resolved contact patterns across the robot body. We therefore study how to integrate such whole-body tactile information into vision-language-action (VLA) policies for contact-rich control. Our approach, Uni-VLaT, introduces a tactile pathway whose latent state is trained not only for action generation, but also to predict future tactile, proprioceptive, and visual representations. This predictive objective builds a tactile-anchored multimodal context, encouraging a more structured understanding of the physical world. We evaluate Uni-VLaT on five real-robot tasks covering tactile-triggered locomotion, sustained physical interaction, human-robot contact, and loco-manipulation. Uni-VLaT achieves a 75% average success rate, outperforming a baseline without tactile input by 43 points and a tactile-input baseline without predictive supervision by 7 points. Across two pretrained VLA backbones, our method improves Table Sweeping by 30 points on both backbones and Back-Tap Walking by 85-90 points. Ablations further show that contextualized tactile prediction and absolute future targets are critical to performance. These results indicate that predictive tactile learning provides an effective route for extending pretrained VLA policies to whole-body physical interaction.

Yuefan Wang, Huaicheng Zhou, Xiao He, ... , Jinxin Liu, Donglin Wang
总结: 提出通用低延迟人形全身遥操作统一框架GAE,实现多样动态全身行为的实时人机同步复现。
方法: 融合视频、动画与动捕构建大规模人类运动数据集,特权生成器策略在仿真中跟踪参考产出可行轨迹,可部署执行器经课程域随机化学习跟踪。
证据: 借助异构数据与两阶段范式,系统在真实人形上展示了低延迟、多样且同步的全身遥操作能力。
为什么适合我: 其生成器-执行器分离与我全身运动跟踪策略设计同构,可借鉴用于清洗噪声人类参考并部署到真机的训练流程。
原摘要

Humanoid avatars extend human physical presence beyond the body, enabling people to participate in social, service, and labor activities through remotely operated robots. This requires teleoperation systems capable of realizing diverse and dynamic whole-body behaviors while maintaining responsive human-robot synchronization. We present General Action Expert(GAE), a unified learning framework for general-purpose, low-latency humanoid whole-body teleoperation. To cover diverse human behaviors, GAE builds a large-scale human motion dataset from heterogeneous sources, including videos, animations, and motion capture, followed by standardization and augmentation. GAE then addresses the noise and embodiment mismatch in human motions with a two-stage training paradigm: a privileged generator policy first tracks human motion references in simulation and rolls out feasible humanoid trajectories; a deployable executor policy then learns to track these generated trajectories under curriculum domain randomization. For responsive human-robot synchronization, GAE introduces a latency-conditioned anticipation mechanism that adaptively compensates for end-to-end delay during real-time teleoperation. Simulation and real-world experiments on Unitree G1 and Westlake O1 robots demonstrate that GAE enables humanoids to smoothly mirror diverse, agile, and expressive human behaviors. Project website: https://wangyf0928.github.io/gae-wlrobotics/

Chuan Qin, Shaoting Zhu, Siyuan Luo, ... , Hongyu Zhao, Hang Zhao
总结: 提出融合显式全身动作监督的生成式视频世界动作模型WB-WAM,为人形移动操作预训练身体与灵巧手异构先验。
方法: 以统一物理动作空间整合身体、根与手部标注,对1880小时部分标注数据做视频-动作联合学习,再经PICO中期训练与前向运动学辅助监督适配机器人。
证据: 仿真HumanoidArena全身任务达81.9%,真实世界实验进一步验证模型在真机上的有效性。
为什么适合我: 展示用大规模人类视频先验补足机器人数据稀缺的路线,可为我的人形全身控制器提供可迁移的表征初始化与任务适配方式。
原摘要

Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A shared physical action space integrates body, root, and dexterous hand annotations from heterogeneous sources, enabling joint video and action learning from 1880.2 hours of partially annotated video and motion data. The resulting priors are refined through PICO mid-training and adapted to robot tasks with auxiliary forward kinematics supervision. We construct WB-Datasets to support these stages with retargeted egocentric human demonstrations and robot trajectories, allowing task-aligned human motion to supplement limited robot data. Evaluations in simulation demonstrate strong whole-body task performance with 81.9% in HumanoidArena, while real-world experiments further validate WB-WAM with 84.0% mean success across five tasks. Moreover, task-aligned PICO mid-training improves downstream task performance while reducing the need for real-robot demonstrations. These results support heterogeneous whole-body pre-training and human motion transfer as a practical route to data-efficient humanoid loco-manipulation.

Nam Hee Kim, Markus Kirjonen, Perttu Hämäläinen
总结: 研究动作后果不可逆的高风险高精度运动控制问题,以计算台球为代表提出精英样本驱动的DRL算法SCOOT。
方法: 在优势加权回归基础上仅用精英样本优化策略,结合专家混合网络按状态切换奖励模式,并加入距离正则稳定训练。
证据: 在计算台球基准上评估,显示对稀疏高奖励动作的捕获优于标准AWR,验证了尖锐峰脊奖励地形下的改进效果。
为什么适合我: 虽非腿式任务,但其处理不可逆高精度动作的精英样本与MoE思想,可启发我的跑酷落点等接触敏感技能的RL训练。
原摘要

Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of high-risk, high-precision motion control problems where actions carry irreversible outcomes, driving sharp peaks and ridges to plague the state-action reward landscape. Using computational pool as a representative example of such problems, we propose and evaluate State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications: 1) Performing policy optimization only using elite samples, allowing the policy to better latch on to the rare high-reward action samples; 2) Utilizing a mixture-of-experts (MoE) policy, to allow switching between reward landscape modes depending on the state; 3) Adding a distance regularization term and a learning curriculum to encourage exploring diverse strategies before adapting to the most advantageous samples. We showcase our features' performance in learning physically-based billiard shots demonstrating high action precision and discovering multiple shot strategies for a given ball configuration.

Tianhao Zang, Shanze Wang, Ziqian Wang, ... , Xingjian Xie, Wei Zhang
总结: 面向敏捷四足导航的未来测距预测方法FutureRay,预测各方向未来可行空间与遭遇风险,显著提升动态避障成功率。
方法: 从深度测距历史与自身运动预测多方向多时刻测距,训练强调近期间隙并惩罚高估可用空间,局部规划器结合制动-反应极限选择速度。
证据: 在60个静态与动态仿真场景配对评估中完成率93.3%,优于当前测距持续法的75.0%与另一基线的80.0%。
为什么适合我: 与我的跑酷感知运动相关,控制对齐的未来感知预测可直接对接固定运动策略,提升非结构化环境动态避障能力。
原摘要

Moving obstacles can block a previously clear route while a quadruped robot executes a motion command. We investigate whether predicting changing clearance improves navigation when motion selection accounts for the robot footprint and the time needed to react and brake. We present FutureRay, which predicts ranges across viewing directions and future times, together with encounter risk, from depth-derived range history and observable robot motion. Training emphasizes near-term clearance and penalizes errors that overstate available space. A local planner queries the same forecast for candidate headings and combines it with current observations to check clearance around the robot footprint. Model-based reaction--braking limits guide speed selection, and the resulting velocity commands are passed to a fixed locomotion policy. In paired evaluations on 60 static and dynamic simulation scenes, FutureRay achieves 93.3% completion, compared with 75.0% for current-range persistence and 80.0% for Cartesian Kalman rollout, with perception, planning, and locomotion held fixed. FutureRay also records fewer collisions than both baselines. Qualitative trials on a physical quadruped show avoidance initiated while an obstacle is approaching the route, followed by renewed goal progress. These results show that joint range and encounter-risk prediction can improve obstacle avoidance without retraining the locomotion policy.

Chengqun Yang, Tengjie Zhu, Liang Xu, ... , Xiaokang Yang, Yichao Yan
总结: 提出单步生成的伴语全身动作系统SocialHumanoid,让人形机器人实时产生与语音同步、具情感表达的身体行为。
方法: 给定语音与情感条件,单次前向生成全身动作窗口并以运动历史条件衔接,在线重定向为机器人参考后由全身控制器跟踪执行。
证据: 系统在真实人形上实现低延迟连续生成与物理执行,验证了情感条件控制与实时性的联合可行性。
为什么适合我: 其单步生成加运动历史衔接、在线重定向到全身控制器的完整流程,可迁移到我的生成式运动策略与全身跟踪管线。
原摘要

Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately $6\times$ faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.