Papers for 2026-09-03

10 papers
Loc X. Nguyen, Avi Deb Raha, Huy Q. Le, ... , Dusit Niyato, Choong Seon Hong
总结: 综述人形机器人部署后自进化机制、安全与评估。
方法: 定义含策略感知记忆的状态元组,进化算子慢环更新。
证据: 分析物理身体如何改变自进化范式与现有研究。
为什么适合我: 支持人形全身运动智能持续适应复杂地形接触。
推荐理由: 直接讨论人形机器人,并提到运动强化学习与具身策略,但核心是部署后自进化综述,不是全身跟踪、跑酷或loco-manipulation方法,只算边缘相关。
原摘要

Humanoid robots are becoming an important part of embodied artificial intelligence, driven by advances in reinforcement learning for locomotion, world models for prediction, and vision-language-action models for general control. However, most of these systems remain static after deployment. A policy is trained offline for a fixed objective and then frozen, even though the tasks, environments, and robot bodies keep drifting over time. An emerging paradigm of self-evolving agents aims to address this problem by allowing systems to improve from their own post-deployment experience. Since most existing studies focus on disembodied software agents, this survey examines how self-evolution changes when an agent has a physical body. We first define self-evolution for humanoids and represent a deployed robot using a state tuple that includes its policy, perception, memory, workflow, and body. This state is updated by an evolution operator in a slow outer loop with a lifelong objective. We then organize the literature into four complementary mechanisms of self-evolution, presented in increasing order of autonomy: self-learning, self-adaptation, self-optimization, and self-generation. Since changes to a humanoid can introduce physical hazards, we treat safety and uncertainty as key design dimensions of the evolution operator, and further formulate admissible evolution as a constraint enforced by a world-model verification gate within a human-oversight envelope. Finally, we present that evaluation should track the robot's evolving trajectory rather than a fixed checkpoint, and we identify the lack of a benchmark designed specifically for self-evolving humanoids. Moreover, we outline open challenges spanning AI algorithms, on-board systems, and governance.

Satvik Sharma, Samrat Sahoo, Huang Huang, ... , Dorsa Sadigh, Jeannette Bohg
总结: 单演示泛化多物体灵巧操作,聚焦局部接触几何。
方法: 接触中心奖励鼓励精确接触,提升sim-to-real。
证据: 真实消融达71%成功,跨形状尺度质量摩擦转移。
为什么适合我: 适用于接触丰富loco-manipulation与感知运动。
推荐理由: 接触几何、人类示范模仿和sim-to-real RL接近接触丰富操作,但是多指灵巧手抓取,不是人形或腿足全身控制与loco-manipulation。
原摘要

Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-scale dexterous robot hand data remains difficult. Learning from human demonstrations has emerged as a scalable alternative to robot teleoperation, providing strong priors on object interaction and contact strategies. Recent sim-to-real RL methods incorporate such priors, but often (i) omit rewards that explicitly incentivize precise contact, yielding weak real-world performance, and/or (ii) generalize poorly to unseen object instances. We propose DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points. Its contact-centric rewards encourage precise contact and improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved. Real-world ablations show that DemoMimic achieves 71% success across 16 objects, four tasks, and two robot-hand embodiments, with the smallest sim-to-real drop compared to baselines.

Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Hyeonseong Jeon, Youngwoon Lee
总结: 递归值学习解决长时程离线目标条件强化学习。
方法: 轨迹分解平衡二叉树,从叶到根精确更新值。
证据: 引导深度降对数,误差积累更慢并发现短路。
为什么适合我: 助力长时程全身控制与模仿学习值估计。
推荐理由: 通用长时程离线目标条件强化学习,未针对人形全身跟踪、腿足运动或物理角色控制。
原摘要

Scaling offline goal-conditioned reinforcement learning (GCRL) to long-horizon tasks is difficult because (1) long-range value learning depends on shorter-range estimates that may still be inaccurate, and (2) max-based value backups can amplify overestimation through repeated propagation. We propose DCRL (Divide-and-Conquer RL), which recursively decomposes each trajectory segment into a balanced binary tree and trains the values from leaves to root. Each parent is therefore updated only after its children, using an exact factorization of the observed route rather than selecting among noisy alternatives. Since this objective learns values along demonstrated routes that are not necessarily optimal, DCRL jointly propagates values across trajectories to discover shorter routes. Thanks to the balanced binary tree, DCRL reduces worst-case bootstrap depth from linear to logarithmic, and this shorter dependency structure empirically corresponds to much slower error accumulation. Across diverse goal-reaching tasks, DCRL substantially outperforms prior flat offline GCRL methods, and on the five most challenging long-horizon OGBench tasks, it improves the best prior average score from 55 to 64, surpassing all flat and hierarchical baselines.

Zekai Jin, Huiguang Wang, Xiaoning Sun, Yi Shao
总结: 安装员在环交互RL实现高精度模块装配。
方法: 离线演示稀疏接管接受奖励,Q-chunking策略。
证据: 非更新热启动稳定离线到在线适应过程。
为什么适合我: 人体专长转真机控制,接触丰富场景相关。
推荐理由: 遥操作示范与接触场景交互强化学习有方法重叠,但是工业化建筑精密装配,不是人形全身运动或腿足跑酷。
原摘要

Industrialized construction imposes stringent precision requirements on robotic assembly of modular components such as prefabricated window units. In tolerance-critical operations, the central bottleneck is not only mechanical clearance but also converting tacit installer expertise into data-efficient autonomy under sparse acceptance feedback, contact variability, and millimeter-scale constraints. We present an installer-in-the-loop interactive reinforcement learning framework that acquires expertise through offline teleoperated demonstrations, sparse event-driven binary takeovers at contact-failure boundaries, and acceptance-aligned terminal rewards, logged under a unified schema for traceable offline-to-online adaptation. A temporally abstract action-sequence policy built on Q-chunking with Flow Q-Learning captures multimodal recovery maneuvers under sparse terminal rewards, while a non-updating warm-start phase stabilizes the offline-to-online transition. The framework is evaluated in MuJoCo across the workflow from suction acquisition through clearance-limited seating, under structured staging and end-to-end randomized placement. Within a defined stress-test regime with 2 mm per-side clearance, bounded pose perturbations, and friction randomization, the pipeline attains 100\% autonomous seating with 12--15 min of cumulative installer supervision over 3.0 h of online training, and reaches the 95\% success milestone in approximately 0.5 h and 1.5 h in the two experiments. We also report wall-clock adaptation time, cumulative takeover minutes, intervention-rate decay, and stage-wise failure attribution to inform supervision budgeting. Ablations isolate the complementary contributions of temporal abstraction, installer intervention, and warm-start value calibration.

Tail-Likelihood Reinforcement Learning

2.0/5 偏低 裁判分 2.0 Candidate 当日相对
Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, ... , Jeff Schneider, Andrea Zanette
总结: 尾似然RL优化奖励上尾覆盖而非平均回报。
方法: 最大化超随机阈值对数概率,加权高回报。
证据: 梯度为Best-of-k混合,仅需修改优势函数。
为什么适合我: 增强跑酷等复杂运动中稀有高回报探索。
推荐理由: 通用尾部似然强化学习,未落到角色控制、运动先验或腿足机器人。
原摘要

Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we consider all of its upper tails: for each reward threshold, how likely is the policy to exceed it? This turns a continuous reward into a family of binary success events. We introduce Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold. Its gradient gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients. TailRL requires only a simple modification to the advantage function, making it compatible with existing reinforcement learning pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward training samples to avoid suboptimal solutions and yields models that benefit more from additional samples at inference time.

Kirin: Animal Motion Generation from In-the-Wild Video

2.0/5 偏低 裁判分 1.0 Candidate 当日相对
Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu, Shangzhe Wu
总结: 从野外视频重建并生成四足动物真实运动。
方法: 大规模视频建AiM3D,视觉文本条件引导生成。
证据: 首个大规模对齐视频文本运动四足数据集。
为什么适合我: 动物运动先验可重定向至腿足全身控制。
推荐理由: 从视频学习四足运动先验,略近运动先验,但是面向动画资产生成,而非物理角色控制、重定向或真机全身跟踪。
原摘要

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area lags far behind human motion research due to the scarcity of high-quality motion data. While human motion can be captured in controlled environments, it is impractical for most animal species, resulting in small, domain-limited datasets that restrict downstream applications such as animation. To address this challenge, we introduce Kirin, a framework that reconstructs motion from video, learns motion priors at scale, and generates realistic motion that can be directly applied to animated assets. Using large collections of in-the-wild animal videos, we reconstruct 3D motion sequences and pair them with captions to create AiM3D, the first large-scale dataset offering aligned video-text-motion tuples for quadruped animals. Building on this dataset, we develop a visual-guided motion generation model that conditions on both text and image to guide the generation of realistic motion across diverse animal species. Finally, by leveraging an off-the-shelf image-to-3D model, we automatically rig and animate 3D meshes using generated motion, producing ready-to-render animated animals. Together, our dataset and framework establish a new foundation for large-scale, text and image conditioned animal motion generation and animation. Project page: https://kirin-ani.github.io/.

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Yuyao Zheng, Haipeng Sun, Junwei Bao, ... , Yang Song, Dejing Dou
总结: 势引导策略优化细化多轮稀疏奖励信用分配。
方法: 从组回报估计状态势,差分得跨轨迹优势。
证据: 失败轨迹中有效动作获更好区分信用。
为什么适合我: 改进长时程loco-manipulation的RL分配。
推荐理由: 面向多轮LLM智能体的策略优化,与人形/腿足运动智能无关。
原摘要

Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit assignment, recent work such as GiGPO introduces step-level advantages for intermediate actions. However, these step-level signals still rely on the final outcome of each individual trajectory. As a result, actions within failed trajectories can remain poorly differentiated, so effective actions can receive the same unfavorable credit as erroneous ones. In this work, we propose Potential-Guided Policy Optimization (PGPO) for multi-turn agentic tasks. PGPO estimates empirical state potentials from anchor-state-group return statistics within each rollout group. It then derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation. This provides finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop show strong overall performance relative to recent group-based RL methods. Further analysis provides evidence that PGPO yields more informative failure-side credit signals with negligible training overhead.

Yicheng Sun, Yueyong Lyu, Yuhan Liu, Yanning Guo, Wei Pan
总结: 物理自适应Koopman MPC稳定不确定航天器姿态。
方法: 四元数解析提升线性模型,梯度更新惯性。
证据: 转为高效QP,显著降低在线计算负担。
为什么适合我: 自适应MPC方法可借鉴全身模型预测控制。
推荐理由: 组合航天器姿态的Koopman MPC,虽用模型预测控制,但与腿足/人形运动无关。
原摘要

This paper proposes a physics-based adaptive Koopman Model Predictive Control (MPC) strategy for combined spacecraft attitude stabilization under inertia uncertainties and active target maneuverability. A novel, quaternion-based Koopman model is constructed from a set of analytical lifting functions derived from the quaternion kinematics, which provides a more compact and physically interpretable linear representation of the nonlinear dynamics compared with the conventional black-box EDMD and higher-dimensional DCM-based model. Leveraging the linear structure of this nominal model, a gradient descent-based update law is employed to efficiently identify time-varying inertial uncertainties from real-time input/output data. By integrating this adaptive linear model into the MPC framework, the optimal control problem reduces to a computationally efficient Quadratic Program (QP), thereby significantly lowering the online computational burden compared to nonlinear adaptive MPC. Recursive feasibility and regional input-to-state stability are formally established through the design of terminal ingredients for the MPC. The effectiveness and superiority of the proposed strategy are validated through comparative simulations of an attitude stabilization task for combined spacecraft in a high-fidelity 3D simulator.

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

1.0/5 偏低 裁判分 0.0 Candidate 当日相对
Zihao Wang, Xi Xiang, Yuwen Sun, ... , Fan Li, Wangmeng Zuo
总结: 深度交错文本图像上下文多模态评估基准。
方法: 逻辑时序空间三域八类共2280题评估。
证据: 评估10个SOTA模型在交错场景表现。
为什么适合我: 感知理解相关但与运动智能关联有限。
推荐理由: 多模态大模型图文交错评测基准,不涉及机器人运动控制。
原摘要

Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench

Menghao Li, Linjie Mu, Yin Wang, ... , Yu Zhang, Fanyi Wang
总结: 置信感知在策略蒸馏缓解视觉预测复合误差。
方法: 教师置信纠正学生,严格到宽松转移控制。
证据: 耦合可靠rollout构建与自适应监督。
为什么适合我: 可提升复杂地形感知运动中的视觉蒸馏。
推荐理由: 视觉语言模型的在策略蒸馏,与全身运动和腿足感知运动无关。
原摘要

Autoregressive vision language models unify heterogeneous perception tasks but are highly susceptible to compounding errors. On-policy distillation (OPD) bridges the training-inference mismatch by training students on their own rollouts. However, unreliable student predictions, especially early in training, can derail the trajectory and degrade the quality of teacher supervision. While recent interleaved distillation methods allow the teacher to verify and replace student tokens, they primarily rely on rigid ranking metrics rather than exact teacher confidence, and they overlook how intervention decisions can inform token-level supervision. To address this, we introduce Confidence-Aware On-Policy Distillation (CA-OPD), a framework that couples reliable rollout construction with adaptive supervision. CA-OPD utilizes teacher confidence to selectively correct unreliable student transitions, gradually transferring rollout control to the student via a strict-to-relaxed schedule. Crucially, CA-OPD aligns knowledge transfer with these intervention decisions: corrected positions receive direct cross-entropy supervision from the teacher's prediction, while retained positions benefit from the teacher's full predictive distribution. Evaluated in a multi-teacher setting for GUI grounding and optical character recognition, CA-OPD substantially improves the Qwen3.5-0.8B baseline across all six target benchmarks, including gains of $9.50$ points on ScreenSpot-Pro and $6.72$ points on OCRBench-v2 English. Controlled studies further show that the gains depend on intervention placement, progressive rollout control, and intervention-aligned supervision, rather than intervention frequency alone.