Papers for 2026-08-02

10 papers
Giovanni D'urso, Kaushik Roy, Nicholas Lawrance, Brendan Tidd
总结: 提出CFNBC离线选数框架,高效修复模仿策略视觉鲁棒性。
方法: 生成配对干净干扰观测,测动作漂移选紧凑修复集。
证据: 从大量演示选响应多样集,提升策略对干扰鲁棒。
为什么适合我: 模仿鲁棒修复可借鉴机器人运动模仿与真实迁移。
原摘要

Visuomotor imitation learning has demonstrated success for manipulation tasks. However, the trained policies remain brittle to visual `nuisances', with even minor task-preserving variations such as lighting, distractions or changes in colour result in heavy degradation of the trained policy's performance. While increasing data diversity can improve robustness, it is unclear which additional demonstrations are informative for a particular trained policy. We propose Counterfactual Nuisance Behaviour Cloning (CFNBC), an offline data-selection framework for targeted robustness repair. Starting from a nominal policy trained on `clean' demonstrations, CFNBC generates paired clean and nuisance observations that preserve the expert action, then measures \emph{action drift}: the change in the policy's predicted action under a nuisance that should not alter the desired behaviour. This provides a policy-specific sensitivity signal for selecting a compact, response-diverse repair set from a larger candidate pool, without requiring rollout success labels or online policy execution. We show in MuJoCo bimanual cube transfer and SimplerEnv cube stacking that action drift correlates with nuisance-induced failure, and that response-guided repair with only $20$--$30$ selected candidates substantially outperforms matched-budget random selection while approaching the performance of much larger random repair budgets. These results support a data-centric view of robustness repair: the most useful data are not necessarily the most numerous, visually diverse, or obviously difficult, but the examples that cover fragile response modes of the current policy.

Henrique Silva, Marcelo A. Santos, Guilherme V. Raffo
总结: 提出混合整数跟踪MPC用于无人机非凸走廊导航。
方法: 嵌入最短路径偏移代价与可行性约束保证收敛。
证据: 仿真生成动态有效轨迹,满足走廊约束并收敛目标。
为什么适合我: UAV规划与腿式人形全身接触控制领域差异大。
原摘要

This work presents a motion planning framework for UAV navigation in non-convex urban air corridors. The planner is based on a mixed-integer tracking model predictive control formulation that enforces corridor feasibility and dynamic consistency within a single optimization problem. To guarantee convergence to the target and mitigate the occurrence of local minima induced by non-convex geometry, a shortest-path-based offset cost with feasibility constraints is embedded directly into the planning problem. Numerical simulations show that the proposed formulation generates dynamically valid trajectories that satisfy the corridor constraints and converge to the target without relying on external global planning stages.

Zahra Abdalla Elashaal, Afef Hfaiedh, Nahla Khraief, Issmail Ellabib, Giansalvo Cirrincione
总结: 提出两层HRL-SAC框架解决稀疏奖励长视野探索。
方法: 高层战略规划,低层SAC连续控制,熵正则优化。
证据: SAR-2上优于平坦SAC,成功率覆盖效率与收敛更好。
为什么适合我: 层次熵正则RL可应用于机器人长视野全身运动控制。
原摘要

Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the continuous-control Soft Actor-Critic (SAC) algorithm, and they utilize entropy-regularized policy optimization. The proposed framework was trained and evaluated using the Search-and-Rescue-2 (SAR-2) dataset. HRL-SAC effectively addresses sparse-reward long-horizon search problems characterized by delayed rewards and continuous control, and its outperforming the flat SAC baseline reinforcement learning in terms of success rates, coverage efficiency, and convergence. These findings indicate that hierarchical entropy-regularized policies are a promising solution to tackle long-horizon sparse-reward reinforcement learning tasks.

Tleukhan Mussin, Yafei Ou, Mahdi Tavakoli
总结: 提出PBD缝合与MPM软组织仿真环境支持手术RL。
方法: 引入接触耦合实现缝合组织双向摩擦拖曳交互。
证据: GPU多流并行优化,支持多场景并提供RL接口。
为什么适合我: 接触丰富软体仿真可借鉴腿式地形接触建模。
原摘要

Recent advances in robotics research have created a strong demand for high-performance simulators. Surgical robotics simulation faces unique challenges due to the need to model diverse objects, such as rigid instruments, soft tissue, and fluids. While many studies simulate sutures or soft tissue independently, only a few have considered the complete soft-tissue suturing scenario, including the contact between sutures and deformable tissue during suture insertion. Building on previous work, this paper presents a novel suturing simulation environment using sutures modelled by position-based dynamics (PBD) and soft bodies modelled by the material point method (MPM) while considering two-way contact with frictional and drag forces. We introduce a contact coupling method between the PBD suture and the MPM soft tissue, enabling visually plausible suture-tissue interactions. The simulator is optimized for GPU execution with parallel scenes using multiple CUDA streams, and we present a Reinforcement Learning (RL) environment for autonomous suturing sub-tasks, including needle insertion, driving, and extraction. Using ML-Agents, RL agents trained in the simulator show stable learning and achieve 80% and 68% success rates in needle insertion and extraction, respectively, under the strictest distance threshold.

Gong Gao, Xiao Lai, Ziqi Xie, ... , Xianhui Liu, Weidong Zhao
总结: 提出CWAC框架缓解离策略连续控制过估计偏差。
方法: 分布评论家建模不确定性,协作加权悲观评论家。
证据: 显式考虑预测不确定性,改善actor-critic训练稳定。
为什么适合我: 离策略RL改进有助于机器人全身连续控制策略学习。
原摘要

Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estimation errors. The recursive propagation of such errors leads to persistent overestimation bias and degraded training stability in actor-critic methods. Existing approaches attempt to alleviate this issue via prioritized sampling or modified value learning objectives, but often overemphasize high-uncertainty transitions caused by limited data coverage or bootstrapping errors, thereby further amplifying bias.In this paper, we propose Collaborative Weighting Actor-Critic (CWAC), a unified framework that explicitly accounts for predictive uncertainty in value estimation. CWAC employs distributional critic to model return uncertainty and introduces a collaborative weighting mechanism that jointly reweights TD-errors and uncertainty, enabling robust learning from reliable samples while suppressing noisy updates. In addition, we incorporate a stochastic pessimistic value estimation scheme via sampling from the return distribution, which effectively mitigates error propagation during policy improvement. CWAC can be seamlessly integrated into existing off-policy algorithm frameworks such as SAC, TD3, and DDPG with minimal overhead. Empirical results demonstrate that our proposed method significantly enhances performance across a diverse range of simulated tasks. Our code is publicly available at https://anonymous.4open.science/r/CWAC-348E.

Wenwu Fan, Qihong Lin, Zhijie Xia, ... , Qiang Chen, Liangsheng Zhu
总结: 提出ACRL自适应控制训练推理差异稳定LLM强化学习。
方法: 自适应维持差异合理范围,增熵增强探索。
证据: FP8量化时维持差异,提升训练稳定与准确率。
为什么适合我: 针对LLM训练,与机器人运动控制仿真迁移无关。
原摘要

Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference discrepancy stems from two primary factors: an architectural separation between training and inference engines, and the use of low-precision quantization in inference versus higher-precision computation in training. To address training instability issues caused by high training-inference discrepancy, we present the principles and methods for its adaptive control. We propose Adaptive Control Reinforcement Learning (ACRL), which adaptively maintains the training-inference discrepancy within a reasonable range to ensure stable RL training. Beyond stabilization, ACRL inherently increases policy entropy, thereby enhancing exploration and improving accuracy. The experimental results show that when the inference engine utilizes FP8 quantization, ACRL consistently maintains the training-inference discrepancy within a reasonable range and stabilizes RL training. Furthermore, ACRL not only matches the accuracy of the BF16 baseline but also outperforms importance sampling (IS) fixes.

Xiangcheng Zhang, Yilun Du
总结: 提出世界动作规划器结合VLM与世界模型可泛化决策。
方法: VLM提初始计划,世界模型滚动优化搜索精炼动作。
证据: 组合任务新布局零样本显著优于VLA与WAM策略。
为什么适合我: 动作条件世界模型规划可应用于人形移动操作泛化。
原摘要

Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific training environments, they often fail to generalize to novel scenes and tasks. In this work, we propose World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world model. Our system enables an agent to propose initial action plans and iteratively refine them via optimization and search, reasoning over imagined world model rollouts. We demonstrate that our approach achieves superior performance across compositional tasks, new layouts, and zero-shot generalization scenarios, significantly outperforming state-of-the-art end-to-end policy models such as VLAs and WAMs. Project website at worldactionplanner.github.io

Emre Özkaya, Nicolas R. Gauger
总结: 提出参数化奖励塑造解决非完整约束停车局部最小。
方法: 覆盖门控对齐反馈,贝叶斯联合优化奖励与超参。
证据: 共优化DQN解决失败模式,成功率与轨迹平滑更优。
为什么适合我: 奖励塑造可参考敏捷运动RL,但任务为汽车停车。
原摘要

Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we show that environmental reward parameters and algorithmic hyperparameters are deeply co-dependent, requiring joint meta-optimization to achieve stable convergence. By employing surrogate-based Bayesian optimization, our co-optimized Deep Q-Network (DQN) agent resolves characteristic control failure modes, significantly outperforming uncalibrated baselines across both success rate and trajectory smoothness.

Anduel Mehmeti, Gabriella Gigante, Salvatore Venticinque
总结: 探索可解释RL辅助空中交通管制高风险决策。
方法: 简化ATC训练避禁飞区智能体,用显著性图解释。
证据: 初步可解释方法有助于建立高风险环境AI信任。
为什么适合我: 航空可解释RL与腿式人形全身运动控制无关。
原摘要

To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviation--and to advance toward higher levels of automation and seamless human-AI collaboration--building trust in AI-driven solutions is essential. Trust, in turn, is closely linked to the explainability of AI systems. The rapid advancements in AI across various domains have underscored the challenges of establishing trust, raising increasing interest in AI explainability even more when applied to deep learning. In this context, the present work aims to explore the application of explainability techniques to Reinforcement Learning (RL) algorithms, specifically within the safety-critical domain of Air Traffic Control (ATC). Using a simplified ATC environment as an initial testbed, an intelligent agent is trained with a reinforcement learning algorithm to make decisions on alternative flight routes that avoid no-fly zones. As a preliminary explainability approach, a saliency map is employed, providing insights into the input features that most significantly influence the agent's decision-making process.

Zhefeng Huang, Yilin Cai, Ankit Patel, ... , Brendan Browne, Yue Chen
总结: 提出FoMo-FD流匹配世界模型检测手术模仿失败。
方法: 学名义视觉动态,评逆传输非一致性检测不一致。
证据: 无需失败演示,共形校准阈值实现窗口级检测。
为什么适合我: 世界模型失败检测可借鉴机器人模仿策略安全部署。
原摘要

Imitation learning has shown increasing promise for autonomous robotic surgery, yet safe deployment remains challenging due to the safety-critical nature of surgical tasks and the complexity and variability of surgical environments. Failure detection is therefore an essential safeguard, but its development remains difficult due to the challenges of scarce failure data, highly variable manipulation dynamics, and the need to balance missed detections against disruptive false alarms. To address these challenges, we introduce FoMo-FD (Flow-Matching World Model for Failure Detection), a failure detection method that learns nominal short-horizon visual dynamics with an action-conditioned flow-matching world model. FoMo-FD scores the inverse-transport nonconformity of observed endpoint latents, enabling window-level detection of visual-action inconsistencies without requiring failure demonstrations. Detection thresholds are obtained by conformal calibration on successful executions, yielding task-specific alarms without assuming future failure types. We evaluate FoMo-FD on four surgically relevant manipulation tasks with twenty failure modes across simulation and real-world experiments using the da Vinci Research Kit (dVRK). Results show that FoMo-FD outperforms observation-level anomaly baselines and a prediction-error variant of the same world model, with the wrist-camera view achieving the strongest performance, including a 96.6% failure detection rate (FDR) at a 1.3% false alarm rate (FAR).