Papers for 2026-09-26

10 papers

Rolling-WAM: World Action Models with Rolling Imagination

2.0/5 偏低 裁判分 4.0 Candidate 当日相对
Yinghua Zhou, Junjie Ye, Yiqi Zhao, ... , Vitor Guizilini, Yue Wang
总结: 滚动世界动作模型分布去噪降低延迟提升闭环响应。
方法: 滑动窗口交错噪声级,逐步全去噪近端动作块。
证据: LIBERO、RoboTwin和Unitree G1上达竞争性操作性能。
为什么适合我: 人形操作适用,契合扩散模型与全身闭环控制。
推荐理由: 操作域的世界动作模型与滚动去噪,虽用扩散式生成,但是桌面/双臂操作而非全身运动先验或腿足控制。
原摘要

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.

Doyoung Kim, Edgar Lee, Hyeonsun Park, ... , Uisu Hwang, Seokhwan Jeong
总结: 扭矩观测对齐实现直接驱动夹爪零样本仿真到真机。
方法: 校准扭矩常数、用差分扭矩并注入实测噪声。
证据: 仿真训练教师学生策略成功部署到多指夹爪。
为什么适合我: 接触丰富抓取对齐,助感知运动仿真到真机。
推荐理由: 直驱夹爪抓取的力矩观测对齐与零样本 sim-to-real,只沾边迁移,不是全身控制或腿足运动。
原摘要

Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.

Mehmet Turan Yardımcı, Yunus Emre Çoğurcu
总结: 不确定性门控噪声抑制流匹配VLA在线微调任务崩溃。
方法: 据新颖性与能力信号重分配探索无需任务标签。
证据: LIBERO-10上固定与学习噪声崩溃而门控无崩溃。
为什么适合我: 在线微调VLA利于持续全身运动智能适应。
推荐理由: LIBERO 上 flow-matching VLA 的在线 RL 微调,属桌面通用操作,与腿足/人形全身运动智能无关。
原摘要

Online reinforcement learning fine-tuning of pretrained flow-matching vision-language-action (VLA) policies promises robots that keep learning after deployment, but continued updates often destroy competence on individual tasks while the aggregate still looks healthy. We study this failure mode, which we call task collapse, under a matched small-compute budget on LIBERO-10 with a 450M-parameter SmolVLA policy trained by PPO with stochastic (SDE) sampling. Three exploration-noise policies differ in one live variable: a fixed noise scale, a ReinFlow-style learned noise network, and an uncertainty-gated controller that redistributes exploration across task streams from task-agnostic novelty and competence signals, without task labels or episode boundaries. Under the pooled definition, fixed noise collapses tasks in two of three seeds and learned noise in every seed measured to iteration 200, while the controller collapses none in any of its three seeds. Measured parameter displacement shows the controller's action expert keeps changing, while its mean applied noise is close to the fixed scale in the available logs. The matched comparison supports the controller's effect on task preservation; the separate contributions of its adaptation across states and over time are not disentangled. A lower fixed scale slows the decline but does not stop it. No arm improves on the behavior-cloning baseline in this budget. Two properties of that regime are measured beside this result, not offered as its cause: following the reference recipe, training runs in bfloat16 with no fp32 master copy, under which 96.02% of the action expert's elements stay bit-identical across three consecutive iterations, and an fp32 master copy at the reference learning rate collapses both arms in a single-seed observation. We release tools measuring per-task collapse under four definitions, rescoring noise and instrument tares.

Teeratham Vitchutripop, Alyssa Quarles, Wenhe Zhang, Richard Xue, Daniel Rakita
总结: 流式深度强化学习可实现机器人自适应持续学习。
方法: 预训练后仅用最新经验流式更新适应未见变化。
证据: 机器人成功适应自身与环境等未预见变化。
为什么适合我: 持续适应复杂地形,契合强化学习全身控制。
推荐理由: 流式深度 RL 用于机器人持续适应的一般分析,没有落到人形、腿足或全身运动跟踪。
原摘要

Over the course of a lifetime, robots may encounter novel scenarios unaccounted for in its original training that result in performance degradation. One common approach to mitigating this issue is to further grow the offline training dataset in hopes of producing a policy robust to these changes. In contrast, biological learning occurs moment-to-moment via a stream of experience, unlike the predominantly batch-based and offline nature of deep learning. Although recent works show the feasibility of stream-based deep reinforcement learning, where updates use only the latest experience, none have shown it to be a viable continual learning framework for adapting robotic policies to unseen changes. In this paper, we present the first analysis of streaming deep reinforcement learning for adaptive continual learning in robotics. In particular, we show that, following an initial pretraining phase, streaming deep RL can enable a robot to successfully adapt to unforeseen changes to itself, its environment, or goals. Our primary experiments within quadruped locomotion demonstrate that a deep neural network robotic policy with certain optimizers and plasticity loss mitigation techniques can successfully leverage domain task knowledge from its pretraining to quickly adapt online to diverse changes via stream learning, outperforming batch-based on-policy methods and improving task success rates by up to 90% over the pretrained policy. Furthermore, we perform additional evaluations on robotic manipulation tasks to determine if our previous observations extend to different robotic morphologies and scenarios. Our results show that the successes observed in quadruped locomotion can be partially realized in manipulation with stability and performance limitations. We conclude with a discussion on the limitations of our work and its implications for the future of continual robot learning.

CrossSafe: Towards Cross-Embodiment Latent Safety Filters

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Ihab Tabbara, Yuxuan Yang, Hussein Sibai
总结: 跨本体潜在安全过滤器共享推理差异化动作实现。
方法: 本体条件安全过滤适配形态运动学与动力学差异。
证据: 安全推理跨机器人共享而动作安全依赖本体。
为什么适合我: 跨人形腿足安全过滤利于loco-manipulation。
推荐理由: 跨本体 VLA 的潜空间安全滤波,关注操作安全而非运动跟踪、跑酷或 loco-manipulation。
原摘要

Cross-embodiment learning has shown that a single model, such as a vision-language-action (VLA) model, can learn state representations and manipulation skills that can be applied across heterogeneous robots to accomplish various tasks. We hypothesize that the same holds for safety enforcement. The reasoning required to satisfy a safety constraint, such as detecting an obstacle, recognizing that it should be avoided, and selecting a safe abstract action, is largely shared across robots. What differs across embodiments is how the abstract safe action is realized: morphology, kinematics, and dynamics determine which actions are safe and feasible. Consequently, the same action can be safe for one robot and unsafe for another. This is especially important for generalist manipulation policies that operate in a common end-effector action space without explicitly capturing how safety depends on the robot's morphology and kinematics. We propose embodiment-conditioned safety filtering, in which a Hamilton-Jacobi reachability-based value function and its corresponding safety-maximizing policy are shared across robots. Using a morphology-aware latent representation of the robot and its environment, we perform Hamilton-Jacobi reachability analysis directly in latent space so that the learned safety concepts can generalize across embodiments while remaining explicitly conditioned on each robot's morphology and kinematics. We evaluate our approach across five bimanual robot embodiments and five manipulation tasks with whole-body collision-avoidance constraints. Our results show that a single policy, jointly trained across five manipulation tasks and four embodiments, exhibits zero-shot generalization to a held-out embodiment, reducing the nominal policy's collision rate. They also show that training using more embodiments improves generalization.

CAMP: Cooperative Arm-Hand Motion Planning in Constrained Spaces

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Ziyuan Wang, Yunlong Shan, Fei Mo, ... , Xin Jiang, Peng Zhou
总结: CAMP实现约束空间高成功高效协作臂手运动规划。
方法: 可行手纤维刻画耦合,分层搜索加局部臂松弛。
证据: 紧凑表示候选轨迹解决高维搜索与非凸碰撞。
为什么适合我: 接触丰富臂手规划可扩展至全身loco-manipulation。
推荐理由: 约束空间中的臂手协同运动规划,偏几何规划,不是学习式全身控制或人体动作重定向。
原摘要

Coordinated arm-hand motion planning is fundamental to dexterous robotic manipulation in complex and constrained environments. A straightforward solution is to decompose the problem into separate arm path planning and hand motion generation; however, this poses a dilemma: decomposition can miss feasible solutions that require coordinated arm-hand adaptation along the path. Alternatively, directly planning in the high-dimensional joint arm-hand configuration space captures such coupling but faces a substantially enlarged search space and nonconvex collision constraints. To characterize this coupling, we formulate feasible hand fibers that capture collision-free hand configurations for each arm configuration. Based on this formulation, we propose CAMP, a high-success and efficient cooperative arm-hand motion planner for constrained environments. CAMP constructs candidate trajectories through layered hand search with local arm relaxation, then compactly represents them using endpoint-preserving via-point movement primitives (VMPs) for coarse-to-fine joint optimization. Across six constrained simulation tasks, CAMP achieves 84.2-98.5% planning success, outperforming alternative planners with competitive efficiency. Ablation studies verify the contributions of arm relaxation, VMP representation, and coarse-to-fine optimization, while real-robot experiments demonstrate CAMP on constrained manipulation tasks. The project website is available at https://camp-armhand.github.io/.

Jiabin Qiu, Zixuan Chen, Hongye Cao, ... , Jing Huo, Yang Gao
总结: 动作判别世界模型提升反事实MPC动作区分能力。
方法: 残差潜动态与动作恢复正则鼓励保留动作信息。
证据: OGBench-Cube硬启动成功率从3.7%到52.0%等多环境提升。
为什么适合我: 世界模型MPC契合全身模型预测控制与接触场景。
推荐理由: 面向反事实 MPC 的动作可区分世界模型,实验是方块操作,不是全身模型预测控制或腿足运动。
原摘要

Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information. Both objectives encourage planning transitions to preserve action information; their auxiliary heads are discarded at test time, leaving MPC unchanged. On OGBench-Cube, AD-WM improves hard-start success from 3.7% to 52.0% over a matched LeWM baseline and improves mean success over the reproduced baseline in four of five simulation environments. Planning diagnostics show that factual prediction error and whole-bank action ranking do not follow the closed-loop success ordering, whereas CEM-aligned elite regret tracks success more closely. With a frozen V-JEPA 2 encoder and matched DROID post-training, AD-WM also improves zero-shot transfer to our Franka setup, increasing basic pick-and-place success from 42.2% to 71.1% without lab-specific adaptation. These results suggest that world models for planning should preserve action-dependent differences needed for counterfactual selection, rather than optimize factual prediction accuracy alone. More videos and code are available at https://ad-wm.github.io/.

Jinchang Zhang, Jiakai Lin, Guoyu Lu
总结: FlyCNS基于果蝇连接组组织信息实现通信受限控制。
方法: 局部传感运动计算加选择性上下行通路通信。
证据: Unitree Go1仿真验证弱先验通信与强化学习策略。
为什么适合我: 腿足分布式全身协调利于复杂地形感知运动。
推荐理由: 以果蝇连接组为先验的通信约束具身控制,虽涉及肢体协调,但不是人形/腿足运动智能或运动先验。
原摘要

Robotic bodies are inherently distributed in sensing and actuation, yet learning-based control still commonly relies on centralized information processing. This work studies the problem of information organization in communication-constrained embodied control: which computations should remain local, and which information is worth transmitting for whole-body coordination. We propose FlyCNS, an embodied information-organization framework inspired by the Drosophila brain--nerve-cord connectome. FlyCNS preserves local sensorimotor computation within each limb and enables selective long-range communication through separate ascending and descending routing pathways. From a real connectome, FlyCNS extracts the directional structural complexity of these two pathway types and uses it as a weak prior over communication allocation, while message content, transmission timing, and locomotion policies remain task-adaptive and are learned through reinforcement learning. In Unitree Go1 simulation, FlyCNS exhibits more graceful performance degradation as the communication budget is tightened. Under the most restrictive setting, it uses only about 21--22\% of the communication of the full-communication reference, while still maintaining a tracking score of approximately 0.882 under both command protocols, with a gap of no more than 6.1\% from the full-communication reference. These results indicate that real neural connectomes can inform not only the structural design of control networks, but also provide transferable inductive biases for information organization across embodiments, guiding robots in balancing local computation and long-range coordination under limited communication resources.

Kyosuke Minomo, Ryo Takahashi, Kotaro Yasui, ... , Motoji Yamamoto, Ayato Kanada
总结: 高关节密度提升蛇形机器人障碍辅助运动性能。
方法: 关节可重定位实现高密度架构并比较高低密度。
证据: 高密度抑制反应力突变避免低密度停滞卡住。
为什么适合我: 接触丰富障碍地形启发腿足复杂地形运动适应。
推荐理由: 蛇形机器人的障碍辅助运动,虽是复杂环境移动,但不是腿足/人形全身运动或 loco-manipulation。
原摘要

Obstacle-aided locomotion is a fundamental capability for snake robots to traverse complex environments. However, conventional rigid-link snake robots often suffer from stagnation or jamming caused by their low joint density (i.e., the number of joints per unit length). This results in discontinuous contact with obstacles, unlike the continuous adaptation of biological snakes. To investigate the effect of joint density on obstacle-aided locomotion performance, we utilized a joint-repositionable snake robot mechanism that decouples actuators from joints, enabling a high-density architecture. We developed two experimental models with identical total lengths but different joint densities (high-density and low-density) and conducted comparative propulsion experiments in obstacle environments with varying obstacle diameters. The experimental results demonstrate that the high-density model substantially suppresses the abrupt shifts in reaction forces that cause stagnation in the low-density model. By maintaining smooth contact points, the high-density configuration reduces power consumption and achieves stable, continuous propulsion. These results highlight high joint density as a key factor in improving the environmental adaptability of snake robots in complex terrains.

Sudip Bhujel, Shanghao Shi, Ruiquan Huang, Ning Zhang, Yang Xiao
总结: TRACE时序梯度反演从策略梯度重建私有轨迹。
方法: 自回归利用跨时相关与动作闭式恢复重建序列。
证据: 达18.8dB PSNR近完美动作恢复且毫秒级重建。
为什么适合我: 具身RL轨迹隐私对全身运动学习数据安全有启示。
推荐理由: 虽涉及具身强化学习中的观测-动作轨迹,但核心是梯度反演隐私攻击,而非人形/腿足全身控制、运动跟踪或 loco-manipulation。
原摘要

Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches $18.8$ dB PSNR with near-perfect action recovery at $3$-$4.5$ ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster. Further evaluation demonstrates TRACE's broader applicability across recurrent, residual, and compact transformer victim architectures, multi-modal inputs, and larger discrete action spaces. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.