Papers for 2026-09-13

10 papers
Yizhan Li, Jianxin You, Mengyang Xiong, ... , Dongqing Zhang, Bang Liu
总结: 首个物理接地基准,测MLLM在人形突发危险中的即时反应决策。
方法: MLLM作模拟人形大脑,用240Hz刚体仿真生成无标注真值场景。
证据: 涵盖17事件族超1000可复现场景,含外观矛盾物理的对抗物体。
为什么适合我: 直接相关人形接触丰富场景反应,助力全身感知运动评估。
推荐理由: 虽用仿真人形面对突发危险,但是多模态大模型反应决策评测,不是全身控制、模仿或运动跟踪方法。
原摘要

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing sudden household hazards; it spans 17 event families and over 1,000 bit-for-bit reproducible scenes with exact, annotation-free ground truth derived from 240 Hz rigid-body simulation, including adversarial objects whose appearance contradicts their physics (a foam anvil, a steel apple). We further design a five-metric suite that scores each reaction along three axes: reasonable, safe, and physically grounded. We physically execute every committed plan so that decisions have observable consequences. With this harness we evaluate seven representative MLLMs. Results show that reactive safety is far from solved: models mishandle roughly one hazard in three, act from fixed dispositions rather than the observed scene, trust appearance over motion, and miss interception points at meter scale even when the chosen action is correct; none of these failures shrink with model scale. ReactHuman thus offers both a fine-grained diagnosis and a scalable training signal toward physically grounded, safety-aware embodied agents. The benchmark can be found here: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled

Muyuan Ma, Yi Zhang, Yang Yang, ... , Yan Yang, Yue Xie
总结: 可变几何桁架结合接触语义原语实现形态计算与控制。
方法: 提取四接触语义原语,物理投影适配新台阶,全空间QP跟踪。
证据: 0.10m至0.075m转移时目标函数评估减少63.7%。
为什么适合我: 接触形态适应启发腿足复杂地形与接触丰富控制。
推荐理由: 可变几何桁架加轮式底盘的接触语义越障,不是腿足或人形全身控制。
原摘要

Reconfigurable robots can change their contact geometry when a fixed body cannot negotiate an obstacle. A variable-geometry truss (VGT) distributes this shape change through a load-bearing structure, but coupling it to a mobile base creates a high-dimensional coordination problem. GeoTrussRover combines an electrically actuated VGT, a wheeled base, and contact-semantic morphology planning and control. We solve one source traversal and extract four contact-semantic primitives that describe coordination among 21 members. Physics-constrained projection adapts them to unseen step heights with the same contact topology. When every phase remains feasible, adaptation does not recompute the complete motion. If one phase violates the new physical constraints, only that phase is recomputed. A full-space QP then tracks the adapted motion and corrects member and wheel errors. For transfer from 0.10m to 0.075m, the method reduces objective-function evaluations by 63.7% relative to full recomputation. Contact-phase feasibility analysis covers step heights from 0.10 to 0.46m, or 1.08 to 4.97 wheel radii, with the upper value near the theoretical feasible boundary. The electric prototype traverses 2.11 wheel radii. The resulting low-dimensional representation stores task coordination in a hyper-redundant, load-bearing morphology and reuses it during locomotion.

Tengbo Yu, Jiahao Wu, Daohan Li, ... , Xiaojian Ma, Hangxin Liu
总结: 人机共享外骨骼实现一对一灵巧接触丰富演示迁移。
方法: 共享编码器与腕相机,重定向成跨本体监督,直接训原图策略。
证据: 五接触任务数据效率高3.0倍,平均成功率70.0%。
为什么适合我: 完美匹配接触丰富模仿重定向,支持人形loco-manipulation。
推荐理由: 虽涉及示教采集与跨本体监督,但核心是可穿戴外骨骼上的灵巧手,属于明确不感兴趣方向,且不是人形全身遥操作。
原摘要

Imitation learning for dexterous hands is bottlenecked by the difficulty of collecting contact-rich demonstrations that transfer faithfully to the robot. Prior wearable-exoskeleton systems record only on the human side and retarget via open-loop mappings calibrated in free space, which degrade under contact. We present SEED-UMI, a framework in which both the human and the robot wear the same exoskeleton: joint encoders become a physically shared measurement, and wrist cameras mounted to the exoskeleton observe the same outer mechanism during both human data collection and robot policy rollouts. This turns retargeting into paired cross-embodiment supervision and lets policies train directly on raw exoskeleton-centric wrist images, without segmentation or inpainting. On five contact-rich tasks, SEED-UMI achieves 3.0x greater data collection efficiency than exoskeleton-based teleoperation and a 70.0% average rollout success rate.

Yanhong Liang, Xianwei Liu, Chaojie Fu, ... , Wei Yang, Hongtao Wang
总结: 图模仿与声学模型使灵巧手高保真演奏复杂钢琴曲目。
方法: RL框架加图优化生成类人手指策略,耦合模型调按键速度。
证据: 定量评估显示高保真表现多样曲目与动态变化。
为什么适合我: 灵巧接触模仿可迁移至人形手部全身控制。
推荐理由: 灵巧手弹琴用 RL 与类人指法模仿,仅边缘接近人体动作模仿;非人形全身、腿足跑酷或 loco-manipulation。
原摘要

Enabling robots to perform musical instruments with human-level expressivity represents a frontier in bridging the gap between mechanical precision and artistic interpretation. Despite advances in robotic dexterity, replicating the fluid finger transitions and nuanced dynamic control characteristic of human pianists remains a significant challenge. Through a reinforcement learning-based control framework, we demonstrate that a dexterous robotic hand can achieve high-fidelity performance across a diverse piano repertoire. Central to our approach is a graph-based optimization strategy that guides the robot to generate natural pre-press and key-press fingering strategies that closely resemble human movement patterns. To achieve expressive sound production, the control system is coupled with a physics-inspired acoustic model that modulates keypress velocity to accurately reproduce the dynamic variations specified in musical scores. Quantitative evaluations demonstrate that our expressive control model significantly outperforms baseline methods in both finger morphology similarity and dynamic velocity accuracy. In a perceptual test involving participants from diverse listener groups, performances generated by our system are significantly preferred over baseline robotic performances and are indistinguishable from human performances for non-professional audiences. Furthermore, extensive experiments across multiple musical styles confirm that our method maintains high note-level accuracy while achieving expressive performance. Our approach provides a robust pathway for robotic systems to move beyond mere mechanical accuracy, elevating robotic musicianship to a level of expressive performance comparable to human pianists.

Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

1.0/5 偏低 裁判分 3.0 Candidate 当日相对
Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet
总结: 多步转移前瞻RL近最优规划,虽NP难但有高效近似。
方法: 随机多项式时间近似方案,扩展处理未知转移。
证据: 风电场存储控制基准验证方法有效性。
为什么适合我: 前瞻RL规划可提升人形复杂场景决策运动智能。
推荐理由: 多步转移前瞻强化学习的复杂度与近似算法,纯理论,非运动控制。
原摘要

We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before deciding its course of action. Although look-ahead can substantially improve achievable performance, [1] showed that optimal planning with multi-step transition look-ahead is NP-hard. However, this hardness was established using a discount factor close to one. It was therefore unknown whether the problem remains hard for every discount factor, and whether near-optimal planning can nevertheless be performed efficiently. We resolve both questions. First, we show that for every fixed discount factor, exact planning remains NP-hard. Second, we introduce a randomized polynomial-time approximation scheme for every fixed look-ahead depth. Third, we extend our approach to account for unknown transitions. We empirically validate the soundness of our results on the wind-farm storage-control benchmark of [2], showing that our approach, optimally accounting for -step look-ahead information, offers substantially better performance than existing algorithms. [1] Corentin Pla, Hugo Richard, Marc Abeille, Nadav Merlis, Vianney Perchet : On the Hardness of Reinforcement Learning with Transition Look-Ahead [2] Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman : Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach

Tien Dat Vu, Minh Doan
总结: 预定义时间弹性积分RL处理未知系统约束攻击扰动。
方法: 积分RL用数据训批评家,Lyapunov证预定义时间收敛。
证据: 两连杆机器人操作器稳定控制验证方法有效。
为什么适合我: 鲁棒数据驱动RL适合腿足约束全身控制。
推荐理由: 未知非线性系统的预定时间积分强化学习与抗攻击控制,与机器人运动智能无关。
原摘要

This paper investigates optimal control for nonlinear systems with unknown dynamics, input constraints, disturbances, and adversarial signals. The objective is to develop a learning-based control method that allows the designer to prescribe the desired convergence time in advance. An integral reinforcement-learning framework is proposed to avoid requiring exact knowledge of the system dynamics while ensuring that the control input always satisfies the actuator constraints. Current and recorded data are combined to train the critic without requiring persistent excitation. The learning gain is selected directly from the prescribed convergence deadline. Lyapunov analysis is then used to establish practical predefined-time convergence of the coupled state-critic system in the presence of disturbances and adversarial channels. The effectiveness of the proposed method is validated through the stabilization control of a two-link robot manipulator.

Adam Haroon, Cody Fleming
总结: 认证安全策展为离线RL提供分布无关训练集保证。
方法: 过滤克隆管道,段比较训价值,校准阈值认证选择。
证据: 策略满足成本预算,无合适阈值时拒绝。
为什么适合我: 安全离线策展保障人形模仿与全身控制安全。
推荐理由: 离线安全强化学习的训练集认证与行为克隆理论,无人形或腿足场景。
原摘要

Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(α, δ)$ bound on the unsafe fraction of the selection, and behavior cloning follows. What is certified is the training set, not the policy, whose cost we report rather than bound. Where no threshold attains the target, the procedure refuses. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on twelve of fifteen DSRL tasks, matching a clone of the ground-truth safe subset, which needs a label on every trajectory. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity of the pool's top quantile, which the calibration sample estimates and through which the scorer enters.

A Bellman Optimality Equation for Plasticity

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Jeremy Lucas, Doina Precup
总结: 提出塑性Bellman最优方程,优化持续RL中塑性。
方法: 在MDP中建立类似赋权的塑性优化Bellman方程。
证据: 初步工作证明存在塑性优化的Bellman最优方程。
为什么适合我: 持续塑性优化有助人形终身运动学习适应。
推荐理由: 持续强化学习中可塑性的贝尔曼最优方程,与腿足/人形控制无关。
原摘要

In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et al. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent's observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Markov decision processes. We show that there exists a Bellman optimality equation for optimizing plasticity similar to previous work for empowerment.

Bin Lei, Yu Li, Prafulla Kumar Choubey, ... , Silvio Savarese, Chien-Sheng Wu
总结: 信念转移分支优化树结构RL分叉获步级信用。
方法: 读答案信念,在连续信念分歧最大处放置分叉。
证据: 三实例无需步标注,提升给定树大小下RL增益。
为什么适合我: 步级信用分配改进可优化全身运动强化学习。
推荐理由: 面向可验证奖励的树结构分叉强化学习,与机器人运动无关。
原摘要

Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \emph{pivots} of the chain's value curve, where the expected outcome turns. We propose \emph{belief-shift branching}: read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most. Three instantiations, none needing step-level supervision, span access levels: a black-box probe, a logit-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training. The signal only \emph{places} forks, and the probe costs about $1\%$ of step compute on mathematics and under $5\%$ on code when it runs inside the rollout engine. In that validation, against Monte-Carlo value curves, a belief-shift signal ranks first in each of the eight model$\times$benchmark panels, ahead of entropy, structural, and LLM-judge baselines. In RL across three model families and two domains, belief-shift forking leads every mathematics aggregate, on OLMo-3-7B by $+2.6$ aggregate and $+2.9$ on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by $+6.5$ on LiveCodeBench-medium.

Tong Li, Saunak Kumar Panda, Yisha Xiang
总结: 对抗状态扰动下风险敏感RL的认证奖励下界。
方法: φ散度松弛扰动,凸优化对偶得可处理认证下界。
证据: 经验法选训练风险厌恶参数改进认证下界。
为什么适合我: 风险敏感鲁棒认证增强人形安全全身控制。
推荐理由: 对抗状态扰动下风险敏感强化学习的下界认证,非腿足或人形控制。
原摘要

Reinforcement learning (RL) agents deployed in real-world environments are often vulnerable to adversarial perturbations in state observations, creating risks in safety-critical applications. Certification methods can improve robustness against adversarial perturbations by providing lower bounds on expected cumulative rewards. Existing certification methods, however, mainly focus on risk-neutral objectives. In this paper, we extend certification methods to risk-sensitive objectives by establishing lower bounds on the exponential utility of cumulative rewards under $l_{p}$-norm-bounded state adversarial perturbations ($1\leq p <\infty$). By introducing a $φ$-divergence relaxation of the perturbation set, we formulate the risk-sensitive certification problem as a convex optimization and derive its dual to obtain a tractable approximation of the certified lower bound. We further propose an empirical method that improves certified lower bounds by selecting the training risk-aversion parameter $β$ independently of the risk level used during evaluation. Experiments on both OpenAI Gym environments and a machine replacement problem show that, compared to risk-neutral training, risk-averse training generally yields policies with higher certified lower bounds, particularly under larger perturbation budgets. Moreover, under both risk-neutral and risk-averse evaluation settings, increasing risk aversion during training leads to non-monotonic certification performance, where certified lower bounds initially improve but eventually decrease due to overly conservative policies.