Papers for 2026-10-02

10 papers

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

5.0/5 很相关 裁判分 10.0 Top pick 当日相对
Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
总结: 测试时进化奖励程序,复用技能解决未见人形loco-manipulation任务。
方法: 对象感知FB基础模型加可调奖励程序,用执行反馈指导规划。
证据: 无需重训即可从尝试中改进并保留所学技能。
为什么适合我: 直接服务人形loco-manipulation与接触丰富全身控制。
推荐理由: 直接做人形 loco-manipulation 的测试时技能重组与接触丰富多阶段任务,最贴近全身运动智能与 loco-manipulation。
原摘要

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

Reactive Humanoid Multi-Contact Using Learned Stability Models

5.0/5 很相关 裁判分 9.0 Top pick 当日相对
Stephen McCrory, Beomyeong Park, Nicholas Kitchel, Nehar Poddar, Robert Griffin
总结: 用学习稳定性模型实现人形多接触反应式稳定。
方法: 采样候选接触并预览质心动力学,学习CoP区域快速评分。
证据: 仿真中冲击韧性平均提升89%,优于无手接触与朴素策略。
为什么适合我: 强化复杂地形接触丰富场景下的全身感知运动稳定。
推荐理由: 人形低稳定场景下的手部多接触支撑与质心动力学规划,紧贴接触丰富全身控制。
原摘要

We present a planning and control approach to reactively use hand contacts to stabilize a humanoid in low stability scenarios, where only using feet contacts may result in a fall. Candidate contacts are sampled within the robot's reachable workspace, and a preview is computed by rolling out the centroidal dynamics through pre-impact, impact and post-impact phases. Sampled points are scored based on the Center of Pressure (CoP) control authority at the post-impact phase. Central to our approach is a learned model of the robot's CoP region during post-impact, which enables rapid evaluation of candidate contact points compared to traditional optimization-based methods. The presented planner has two stages: the first selects an optimal bracing region and the second computes an optimal bracing point within the region. Our simulation results demonstrate an average increase in impulse resilience of 89% over recovery without hand contacts and 17% over a naive planning strategy (closest reachable region). We validate our framework on hardware, performing push tests while standing and walking. The standing trials show an average 43% reduction in stabilization time compared to naive hand placement and the walking trials demonstrate a 18% reduction compared to baseline recovery (without hand contacts).

Kyungmin Lee, Sibeen Kim, Dongyoon Hwang, ... , Jaegul Choo, Hojoon Lee
总结: 多参考RL加速灵巧操作数据的物理可行重定向。
方法: 将重定向建模为多参考跟踪,用离策略RL与几何监督联合学习。
证据: 50动作基准上重定向90%演示,约耗30GPU小时。
为什么适合我: 支持人体动作到真机可跟踪可重定向的全身控制。
推荐理由: 把人手物体演示用多参考强化学习重定向到灵巧手,直接对应人体动作重定向与模仿。
原摘要

Human hand-object demonstrations provide a scalable source of data for dexterous robot learning, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based methods typically optimize each demonstration independently, leading to either limited success under finite simulation budgets or training costs that grow with dataset size. We introduce FlashDexRetarget, an RL framework for multi-reference dexterous retargeting. We formulate retargeting as multi-reference tracking, jointly learning a single policy across many demonstrations with off-policy RL and geometric supervision of the demonstrated interactions. This shared training formulation amortizes optimization across references while enabling the policy to track diverse hand-object interactions. On a 50-motion benchmark from TACO, OakInk2, and HOT3D using XHand and Sharpa Wave Hand as target embodiments, FlashDexRetarget retargets 90% of demonstrations using about 30 GPU-hours, compared with about 46% at about 3,000 GPU-hours for CHORD. This corresponds to about 100 times lower training compute and a 44-percentage-point improvement in retargeting success. Ablations examine the key design choices, while experiments with up to 1,000 motions and real-world replay further demonstrate the scalability and practical applicability of our method.

World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories

3.0/5 一般 裁判分 4.0 Candidate 当日相对
Jiahui Lei, Qianqian Wang, Trevor Darrell, Angjoo Kanazawa
总结: 用稀疏SE(3)轨迹灵活序列建模世界运动。
方法: 流匹配加每令牌噪声与上下文令牌,实现任意边际条件。
证据: 统一人体、手物交互、机器人状态动作等动态实体。
为什么适合我: 为人体动作重定向与全身运动先验提供生成基础。
推荐理由: SE(3) 轨迹生成先验覆盖人体、手物交互和机器人状态,只与运动先验弱相关,并非可上真机的全身跟踪或 loco-manipulation。
原摘要

Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.

Samuel Zhen, Siwon Jo, Yanze Zhang, Wenhao Luo
总结: 全身与附着几何安全框架保障VLA操作避碰。
方法: 构建抓取条件安全集并转为可微CBF约束修改动作。
证据: 在SafeLIBERO基准上实现机器人场景附着几何避碰。
为什么适合我: 提升loco-manipulation中接触丰富场景的全身安全控制。
推荐理由: 面向通用 VLA 操作的全身与附着物碰撞安全,不是人形/腿足运动或 loco-manipulation 控制。
原摘要

Vision-language-action (VLA) policies have demonstrated impressive capabilities in generalizable robotic manipulation, but their deployment in the real world remains challenging due to potential collisions involving different parts of the robot, manipulated objects, and the surrounding environment. Existing inference-time VLA safety frameworks typically rely on simplified end-effector-centered representations that do not explicitly model the full articulated robot and attached-object geometry. In this paper, we present WBAG, a safety framework that models the robot's whole-body and grasp-dependent attached geometry. WBAG constructs a grasp-conditioned safe set that adapts the protected geometry as objects are grasped, then converts this evolving geometry into differentiable CBF constraints that minimally modify the VLA's native six-dimensional operational-space action for collision avoidance across robot, scene, and attached geometry. On the SafeLIBERO benchmark, a variant of LIBERO augmented with obstacles for safety evaluation, WBAG achieves the best overall safety and safe task success among the evaluated methods under a scene-level safety evaluator that monitors all eligible non-task objects, reaching 97.38\% aggregate Scene Safety and 59.38\% Safe Success.

Dehao Huang, Jianbang Liu, Jianpan Gao, ... , Yue Wang, Hong Zhang
总结: 动作相关令牌路由提升VLA强化学习效率。
方法: 学习路由令牌跨层聚合任务特定动作相关特征。
证据: 构建有效状态表示以改进下游动作精炼与价值估计。
为什么适合我: 打通强化学习与模仿学习用于全身操作策略适应。
推荐理由: 通用 VLA 的在线强化学习表征路由,缺少人形全身跟踪、腿足运动或动作重定向。
原摘要

Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.

Xuehui Yu, Eason Yu, Meiyi Wang, ... , Stefano V. Albrecht, Harold Soh
总结: 递归动作相关记忆解决长时程VLA历史依赖。
方法: 优化条件互信息学习记忆函数,分治选择与压缩。
证据: 在长上下文下保持轻量计算并提升历史依赖任务。
为什么适合我: 支持复杂多阶段loco-manipulation的感知运动记忆。
推荐理由: 长时程 VLA 的动作相关记忆,属于通用操作策略,不涉及人形或腿足全身控制。
原摘要

Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information $I(a_t; m_t \mid o_t)$ between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function $m_t = M(h_t)$ and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-$K$ selection over $2K$ tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/

Siddeshwar Raghavan, Ziqin Yuan, Fengqing Zhu, Byung-Cheol Min
总结: 语言扰动基准检验持续模仿学习是否保持接地。
方法: 构建意义保持与改变指令变体,分离能力与语言敏感。
证据: 诊断语义鲁棒性、目标适应与语言敏感性指标。
为什么适合我: 保障人体动作到真机控制中语言引导行为的可靠性。
推荐理由: LIBERO 上的持续模仿与语言接地评测,不是人体动作重定向或人形运动跟踪。
原摘要

Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills. However, retaining task performance does not ensure the behavior remains grounded in language because policies may rely on scene cues, object associations, or memorized task structure. We introduce a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks. We construct meaning-preserving and meaning-changing instruction variants for the Goal, Spatial, Object, and Long suites of LIBERO. Policy experiments focus on LIBERO-Goal, evaluating Original and Paraphrase instructions after each continual-learning stage. We compare representative continual imitation learning methods under their original assumptions while separating task competence from language sensitivity. The proposed diagnostics complement standard learning and forgetting metrics by measuring semantic robustness, goal adaptation, and language sensitivity. Results show that strong continual-learning performance does not always translate to reliable language grounding, and our diagnostics help determine whether retained skills remain correctly guided by their instructions. Additional materials are available at https://sites.google.com/view/stillgrounded

FutureWorlds: Learning Robotic World Models from Alternative Futures

2.0/5 偏低 裁判分 4.0 Candidate 当日相对
Hao Wu, Shengju Qian, Weiyan Wang, ... , Qingsong Wen, Yuxuan Liang
总结: 从替代未来学习机器人世界模型。
方法: 多样束搜索构建候选,有界记忆维护历史,MemSPO优化。
证据: 将视频轨迹奖励转为组相对优势以优化世界模型。
为什么适合我: 为跑酷与loco-manipulation提供动作结果预测先验。
推荐理由: 通用机器人世界模型与候选未来学习,不针对人形全身控制或复杂地形感知运动。
原摘要

Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: https://github.com/Alexander-wu/FutureWorlds.

FERPO: Forward Entropy-Regularized Policy Optimization

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
总结: 前向熵正则策略优化避免批评家动作梯度。
方法: 从熵与KL正则目标导出目标分布,用SNIS拟合演员。
证据: 限制目标偏离滚动策略以提升连续控制策略更新。
为什么适合我: 强化强化学习在全身运动智能中的稳定策略改进。
推荐理由: 通用连续控制最大熵策略优化,仅方法层与强化学习相邻,无具体人形或腿足任务。
原摘要

Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).