Papers for 2026-10-05

10 papers
Jian Hu, Shujing He, Leixin Chang, ... , Chaoyang Shi, Chengzhi Hu
总结: 提出ColoACT系统,用多线索动作分块Transformer策略实现自驱动内窥镜机器人的平滑自主结肠导航。
方法: 在RGB基础上融合估计相对深度与梯度伪高程图,输入ACT类分块策略输出平滑动作块,应对弱纹理、镜面视觉与粘弹性接触。
证据: 论文针对可变形解剖等难点设计并声称实现平滑自主导航,但摘要未给出量化性能指标。
为什么适合我: 接触丰富柔体环境中的自主导航与sim-to-real经验,可借鉴到腿式机器人非结构化地形的感知运动控制。
原摘要

Autonomous colonoscopic navigation can reduce operator burden and the risk of loop formation or tissue trauma, but remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. Existing methods either rely on geometry-driven pipelines, which are efficient and interpretable yet brittle due to manually engineered features and switching logic, or adopt learning-based policies, whose inferred depth/geometry can become temporally inconsistent or overly smooth under weak texture and specular highlights while simulation-trained variants (e.g., deep reinforcement learning) may further suffer from a sim-to-real gap. We propose ColoACT, an autonomous navigation system that integrates an RGB-D-E based Action Chunking Transformer policy (ColoACT policy) for a compact self-propelled Bevel-Gear-Based Endoscopic Robot (BGER). The ColoACT policy augments RGB with estimated relative depth and a gradient-based pseudo-elevation map to enhance fold-ridge saliency and other high-frequency geometric cues, and enables smooth continuous control of the BGER by predicting overlapping action chunks and fusing them via temporal ensembling. In different \textit{ex-vivo} porcine colons (approximately 60 cm), our system achieves success rates of 85.4\% and 72.5\% in straight and curved segments, respectively, and achieves 70\% success in 90-degree turns and 60\% in double-bend sequences, with feasibility further demonstrated in challenging triple-bend segments. The project page is available at: https://Adamhu1.github.io/ColoACT/.

Seonvin Cho, Soohyun Choi, Songnam Hong
总结: 提出RAPO,在离线强化学习中按策略在critic自举与动作执行中的不同角色自适应调整正则化强度。
方法: 通过对候选策略更新求导学习各角色系数,将TD3+BC拆为自举与执行双actor独立调权,IQL则只调逆Q系数并保留原值更新。
证据: 在TD3+BC与IQL上分离两类目标并自适应调权,声称优于共享策略的耦合方式,摘要未列具体基准数字。
为什么适合我: 离线数据上稳定训练策略actor的正则思路,可迁移到真实机器人轨迹上的全身控制策略训练。
原摘要

Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.

Kyochul Jang, Seohyeon Park, Ohchul Kwon, ... , Jongmin Park, Youngjae Yu
总结: 提出人形工具使用基准HumanoidToolBench,联合评测从工具选择到移动执行的完整能力链。
方法: 构建18个任务、覆盖3场景3执行级别与两种工具集的基准,配套ToolBook数据集,含仿真与真机Unitree G1上采集的3.1k条演示。
证据: 评测7个仿真策略与3个真机策略,发现选对工具与完成任务的巨大差距;GR00T N1.7对未见工具选择变差且受无关指令干扰。
为什么适合我: 直接对应人形移动操作与操作-移动协调,是全身控制器上层任务评测与真机数据的重要来源。
原摘要

As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.

Zijie Zhao, Shengqian Chen, Xiaoxu Wang, ... , Yuanheng Zhu, Dongbin Zhao
总结: 提出LocoWM,用世界模型引导的预动残差适配实现兼顾指令跟踪与任务相关物理状态精确调节的高精度运动。
方法: 基础策略负责行走,动作条件世界模型由本体感知历史与建议动作预测未来状态序列,残差适配器据此提前修正;两阶段训练先学运动与动力学再冻结训练适配器。
证据: 声称残差在偏差可观测前即补偿预期偏离,实现高精度运动控制,摘要未给出具体实验数字。
为什么适合我: 预测式残差修正范式可强化足式与人形控制器对地形接触和扰动的高精度响应,值得直接借鉴。
原摘要

High-precision locomotion combines motion-command tracking with precise regulation of task-relevant physical states, enabling robots to interact reliably with their surroundings during motion. Joint end-to-end optimization can leave precision objectives insufficiently optimized, while reactive residual control adjusts actions only after deviations become observable. We present \textbf{LocoWM}, a world-model-guided preactive residual adaptation framework for high-precision locomotion. A base policy provides command-following locomotion, while an action-conditioned world model predicts a sequence of future physical states from proprioceptive history and the proposed base action. A residual adapter conditions on this predicted sequence to generate additive action corrections that compensate for anticipated deviations. Two-stage training first learns locomotion and action-conditioned dynamics, then freezes both modules while training the adapter, separating locomotion acquisition from precision adaptation. Experiments spanning terrain leveling, acceleration compensation, and push recovery demonstrate improved control precision and disturbance robustness over end-to-end and reactive residual baselines. Demos and code are available at: https://zhaozijie2022.github.io/LocoWM

Sanghyuk Park, Kwanwoo Lee, Taekyung Kim, Seohyeon Lim, Yisoo Lee
总结: 提出DODGER,一种在多个动态障碍间导航的安全引导强化学习框架,兼顾训练探索与避碰塑形。
方法: 训练时直接执行策略生成的动作以保持探索,同时用CBF滤波参考与约束违反作为塑形信号,将避碰行为内化进策略。
证据: 经Dubins车安全性分析验证,并在全阶人形仿真与真实机器人上展示多动态障碍间的目标导航。
为什么适合我: 以安全信号塑形而非硬过滤动作的做法,可借鉴到人形在人流密集非结构化环境中的安全运动。
原摘要

Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such safety information into learned policies. However, executing only safety-filtered actions during training can restrict policy exploration, a limitation that becomes particularly consequential in dynamic scenes where safety depends on relative robot-obstacle motion. We propose DODGER, a safety-guided RL framework that directly executes policy-generated actions to drive training rollouts while using CBF-filtered references and constraint violations to shape the policy toward collision-avoidance behavior. We evaluate DODGER through a Dubins-car safety analysis and demonstrate goal-directed navigation among multiple dynamic obstacles in full-order humanoid simulation and real-world humanoid experiments using LiDAR-based perception, without a runtime safety filter.

Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, ... , Jean Oh, Reid Simmons
总结: 提出闭环框架,用LLM对策略rollout的语义化分析迭代修正模仿学习中的结构化策略,无需人工指令。
方法: 把rollout记录为语义表格数据,提示LLM生成诊断分析代码定位策略结构的次优之处并自动迭代改写。
证据: 赛车与开门任务上比零样本LLM结构最高提升15%性能,达到同等强化学习性能所需算力减少75%。
为什么适合我: 自动化失败诊断与策略改写循环,可用于迭代改进全身运动跟踪策略在复杂地形接触下的结构设计。
原摘要

Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.

Siwei Ju, Lu Liu, Jan Peters, Oleg Arenz
总结: 提出可行性感知策略选择VAPS,让人形特技在偏离参考时安全中止或受控跌落以保护硬件。
方法: 训练可随时中止并双脚落地的abort策略与最小损伤跌落策略,学习预测器逐步估计各策略短时域可行性,按最小牺牲层级选择仍可行中最进取的行为。
证据: 随机扰动仿真中大幅减少头部与手部接触等主要硬件损伤来源,同时保持任务完成度。
为什么适合我: 为全身运动跟踪类敏捷动作提供安全切换机制,可显著降低真机部署跑酷等高危动作的硬件风险。
原摘要

Dynamic humanoid motions such as flips risk hardware damage due to suboptimal policies, disturbances or sim-to-real gaps. A motion tracking policy offers no way out once the maneuver leaves its reference, and a backup policy needs to take over to protect the hardware for a minimum-damage landing. Which backup to use matters as much as when to switch. We present Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision. Besides a protective fall policy, we also train an abort policy which can abort the motion at any time, landing on its feet. At every control step, learned predictors estimate whether the nominal tracking policy and the abort policy remain viable over a short horizon, and a least-sacrificial hierarchy keeps the most task-ambitious behavior that remains viable. In simulation with randomized disturbances, VAPS sharply reduces head contact and hand contact, which are the dominant sources of hardware damage, with both a Unitree G1 and a LimX Oli; on the LimX Oli, we validate the viability predictors and the full VAPS controller for side-flip motions. VAPS Pareto-dominates the strongest single-network alternatives we could train, including an end-to-end safe-tracking policy and students distilled from VAPS's own oracle-routed decisions, in both task success and head impact. We also show that VAPS is a powerful framework to supervise undertrained policies and protect the hardware.

Adam Polevoy, Dillon Capalongo, Katherine Tang, ... , Marin Kobilarov, Joseph Moore
总结: 将随机非线性模型预测控制与强化学习结合,实现未知环境中带概率碰撞安全保证的感知导航。
方法: 先用RL训练概率actor-critic与传感器预测模型,再嵌入PAC-NMPC采样框架,以硬约束给出有限时域碰撞概率与值函数下降的统计保证。
证据: 仿真显示该方法提升感知RL导航策略的安全性,并可扩展到高维系统与大传感器输入空间的复杂设置。
为什么适合我: 兼顾RL长期性能与概率安全证书的架构,可为人形在感知受限复杂地形行走时提供安全保障层。
原摘要

In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models. We then leverage these probabilistic models in a sampling-based SNMPC framework known as Probably Approximately Correct (PAC)-NMPC, which uses hard constraints to enforce finite-time statistical guarantees on the probability of collision and value function improvement. By ensuring that our finite-horizon SNMPC policies decrease the value function in expectation, we can approach the long-horizon performance of the RL approach while satisfying probabilistic safety constraints. Through simulation experiments, we show that our approach can improve the safety of perception-based RL navigation policies and scale to high dimensional systems with large sensor input spaces and complex nonlinear dynamics. We also demonstrate our approach through hardware experiments, showing improved performance for vision-based navigation with an agile fixed-wing aerial vehicle in unknown environments.

Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana
总结: 系统综述最优传输在强化学习中的应用,作为弱重叠分布下常用散度度量失效时的替代方案。
方法: 按OT扮演的角色、所比较的分布、OT形式与时间结构处理方式四要素对现有方法分类,并讨论选择动机与实践考量。
证据: 覆盖模仿学习、离线RL与分布偏移部署中用OT对齐策略、专家、数据与模型分布的代表性工作。
为什么适合我: OT距离能对齐参考与生成动作分布,可为运动模仿和对抗运动先验提供比KL散度更稳健的度量工具。
原摘要

Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.

Hangong Chen, Linfeng Cheng, Tahsin Zaman Jilan, ... , Jiaye Wu, Yantian Zha
总结: 提出Video2SwimFish流水线,从真实鱼视频自动重建可控鱼体资产并学习个体化游泳策略。
方法: 多视角视频重建度量尺度可变形网格,VLM actor-critic闭环生成匹配个体形态的内部关节,从中轴曲率提取受真实运动约束的生物运动流形作为低维动作空间。
证据: 发布120尾、6物种的双视角同步数据集及对应可控资产与个体游泳策略,支持带运动真实性约束的学习评测。
为什么适合我: 从视频自动构建可控资产并提取低维运动流形的流程,与运动重定向和物理角色动作模仿高度相通。
原摘要

We present Video2SwimFish, an automated pipeline and benchmark for building controllable fish assets from real-fish videos for underwater embodied AI. Given synchronized multi-view videos of an individual fish, the pipeline reconstructs a metrically scaled deformable mesh from a VLM-selected canonical frame, generates internal articulation adapted to that individual's morphology through a VLM actor-critic loop, and extracts a Biological Locomotion Manifold (BLM) from the fish's observed midline curvature. The BLM provides a low-dimensional action space bounded by real-fish motion, enabling an individual swimming policy to be learned for each reconstructed fish. We release two paired datasets: synchronized top- and front-view recordings of 120 individual fish across 6 species, and the controllable assets and individual swimming policies derived from them. Because every asset is tied to the animal it came from, the dataset supports a benchmark that evaluates locomotion learning not only on task success but on fidelity to that individual in trajectory shape, body curvature, and tail-beat frequency, across trajectory following, reward-free swimming behavior transfer from video, and a downstream case study in which a simulated BlueROV underwater robot captures one of the assets. We find that task success and locomotion fidelity do not necessarily improve together: the method achieving the highest task completion is not the method achieving the highest locomotion fidelity, and we identify faithful reproduction of individual animal locomotion as an open challenge for the community. Project website: https://hangongchen.github.io/video2swimfish-web/.