Papers for 2026-08-17

10 papers
Mao Jiayang, Wang Lanfeng, Peng Zhao-Han
总结: 提出解耦的选项引导残差多智能体强化学习框架用于异质无人艇港口追捕。
方法: 整合共享逃逸信念、角色条件选项目标、自适应规则惩罚与残差策略学习。
证据: 抽象港口场景中OGR-MASAC达75.0%捕获率并具最佳异质协调。
为什么适合我: 水面艇协同追捕与腿式人形全身接触运动控制无关。
原摘要

Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.

Yongyan Cao
总结: 开发预览式相对运动控制器以在脉动组织中调节神经线插入工具尖端。
方法: 估计延迟周期表面运动并短时预测,用无偏置MPC调节相对放置。
证据: MuJoCo中1自由度接触RMS误差1.9微米,优于延迟阻抗与实验室PD。
为什么适合我: 医疗插入工具控制,非腿式机器人非结构化环境运动。
原摘要

Robotic neural-thread placement requires regulating the insertion-tool tip relative to tissue that moves with cardiac and respiratory pulsation. This paper develops a preview-based relative-motion controller that estimates latency-delayed periodic surface motion, predicts it over a short horizon, and uses offset-free model predictive control to regulate relative placement while limiting actuator effort and lateral relative velocity. In MuJoCo, the 1-DOF controller achieves 12.0\um\ free-space and 1.9\um\ contact RMS relative-placement error, versus 18.3/176.8\um\ for delayed-feedback impedance and 286.1/275.5\um\ for lab-frame PD, at the cost of higher peak contact force (3.43 versus 2.00~mN) since offset-free tracking drives the tip fully to the commanded depth rather than yielding against the tissue. In 3 DOF, coupled preview reduces contact lateral shear from 1.34 to 0.50~mm/s with 2.1\um\ lateral RMS error. A feasibility-restored octagonal shear formulation keeps the QP solvable under degraded sensing by adding a bounded shared slack: at 10\um\ RMS per-axis sensing noise, where a matched cost-only controller violates the 0.80~mm/s budget in all 10 seeds (mean/maximum 0.988/1.175~mm/s), the soft-octagon controller completes all 10 seeds with no fallback and no measured violation (0.653/0.712~mm/s), with its operating envelope characterized up to 15\um\ RMS. A two-vertex Lyapunov certificate for the controller's actual finite-horizon error-feedback gain holds over $-40\%/{+}50\%$ reflected-mass mismatch. The modeled tip is a rigid contact point, and the study is simulation-only: flexible-thread and carrier-needle mechanics, a validated transient-force constraint, biological damage thresholds, and hardware-realistic sensing and timing remain required before deployment.

Tran Le Vu
总结: 提出情境质量多样性进化强化学习控制器用于热带商业建筑暖通监督控制。
方法: 维护情境与行为描述符索引的策略档案,共享回放缓冲的进化与SAC算子。
证据: 在新加坡建筑两层简化环境训练并全年回测对比ASHRAE基线。
为什么适合我: 建筑暖通控制与腿式人形机器人运动学习无关。
原摘要

This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical, water-cooled chiller plant and its associated air side. Rather than converging to a single scalarised policy, the controller maintains a product archive of specialised policies indexed jointly by a data- driven operating context, a cluster of daily weather and load regime, and a context-invariant behaviour descriptor, filled by a gradient-free evolutionary operator and a soft-actor-critic policy-gradient operator that share one replay buffer. Every action is filtered through a deterministic safety shield before execution. The controller is trained on a two-tier reduced-order environment representing the latent load, cooling-tower approach and humidity constraints of a Singapore commercial building, and is evaluated over a full annual backtest against an ASHRAE Guideline 36 baseline.

Zetao Hong, Song Yuan, Yuanhao Ding, ... , Zhibin Wang, Chen Tian
总结: 提出MISA-T混合强化学习推出调度准入策略以处理异质会话竞争。
方法: 结合自适应会话准入、工作负载感知KV容量分配与驻留时间感知记账。
证据: 摘要侧重异质推出需求而未报告具体性能数值。
为什么适合我: 大模型推出服务调度,与机器人全身控制器无关。
原摘要

Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with verifiable rewards (RLVR), reinforcement learning from human feedback (RLHF), and agentic rollouts share an asynchronous inference service, their distinct sequence structures, interaction patterns, and KV-residency times create substantially different serving demands. Rollout scheduling must account for this heterogeneity without distorting the workload mixture specified by the trainer. We present MISA-T, a routing-layer admission policy for mixed rollout serving. MISA-T combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting. In rollout-only ablations on Step3.7 and Qwen3.6-35B-A3B, MISA-T improves rollout throughput over a sweep-tuned cache-aware vLLM Router by 53.3% and 43.6%, respectively, while maintaining high prefix-cache hit rates. In a matched 50-iteration Step3.7 experiment, it increases rollout throughput by 35.6% and reduces mean iteration time by 22.8%, while keeping the consumed workload mixture close to the trainer target and achieving comparable task scores.

Zhixin Zhang, Xinke Jiang, Zhibang Yang, ... , Junfeng Zhao, Yasha Wang
总结: 提出LoongReflect将反思建模为记忆控制策略以提升搜索智能体长程能力。
方法: 智能体在可逆轨迹树上使用显式反思与回溯动作巩固事实与证据。
证据: 摘要聚焦局部全局监督不匹配而未给出具体实验数字。
为什么适合我: 语言模型智能体反思,与腿式运动控制与迁移无关。
原摘要

Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue, revise, or abandon the current branch. Learning effective reflection, however, is challenging because reflection is performed locally within the current branch, whereas its utility can only be determined by its contribution to the final trajectory outcome. This local-global mismatch makes outcome-based reinforcement learning provide only local, sparse and delayed supervision for reflective decisions. To solve these, we propose LoongReflect, a training framework that formulates reflection as a memory-control policy. The agent operates over a reversible trajectory tree using explicit reflect and backtrack actions. Reflection consolidates verified facts, missing evidence, and branch-specific risks into working memory, while backtracking removes an unreliable branch from the active context and preserves a concise corrective lesson. To learn this policy, LoongReflect combines two complementary signals through a look-ahead, extragradient-style coordination mechanism. A fast channel distills globally informed reflective behavior from a privileged teacher, with supervision restricted to reflection and backtracking tokens. A slow channel optimizes complete trajectories using outcome-based GRPO, aligning local control decisions with final task success. Experiments on multi-hop retrieval-augmented generation and mathematical reasoning benchmarks demonstrate consistent improvements over outcome-only reinforcement learning and self-distillation baselines.

Zijian Zhao, Sen Li
总结: 证明多智能体独立后继特征组合可产生严格劣于库中所有策略的联合行为。
方法: 分析队友重组改变环境使单智能体价值保证失效的失败模式。
证据: 理论证明独立组合可严格劣于库中每一策略且无单智能体对应。
为什么适合我: 多智能体策略组合理论,与全身运动跟踪关联较弱。
原摘要

Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive. For a single agent, this problem is well understood: successor features with generalized policy improvement, together with their universal extension, recombine a library of learned policies into a policy for any new objective, with a guarantee that the result is never worse than any policy in the library. However, multi-agent transfer has received far less attention, and the common practice of letting each agent recombine its own library independently inherits the recipe but not the guarantee. We prove that this independent composition can produce joint behavior strictly worse than every policy in the library, because recombining teammates changes the environment each agent faces and invalidates the values it relies on, a failure with no single-agent counterpart. We further show that the only unconditionally safe fixed rule is synchronized composition, which moves the whole team to one jointly trained policy but cannot serve objectives that assign different goals to different agents. To attain safety and flexibility at once, we propose MA-USFA, a hierarchical method with two layers: a lower layer of universal successor feature approximators that predicts each agent's successor features while conditioned on its teammates' objectives, and an upper composer that selects, across agents, which library entry each agent should follow and supplies the cross-agent correction a per-agent value cannot represent. Trained once over the distribution of objectives, it is applied at deployment with no per-task adaptation.

Xincong Hu, Lei Ou, Maosen Li, ... , Liguo Hou, Zongzhang Zhang
总结: 提出威胁引导策略感知场景扰动以提升在线强化学习自动驾驶安全性。
方法: 引入策略感知场景编码器捕捉策略行为与周围交互并生成针对性扰动。
证据: 摘要强调针对策略弱点生成场景而未报告具体性能数字。
为什么适合我: 自动驾驶安全场景生成,非腿式人形接触丰富控制。
原摘要

Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.

Rajmund Nagy, Silvia Arellano García, Hendric Voss, ... , Youngwoo Yoon, Gustav Eje Henter
总结: 报告GENEA挑战2026语音驱动手势生成系统的大规模解耦人类评估结果。
方法: 用解耦评估分离运动质量与语音对齐,并引入语义手势任务与文本错配。
证据: 四项研究收集超23000票,数据集片段以68-95%胜率显著更高。
为什么适合我: 手势运动生成评估,与物理仿真角色运动模仿有轻微关联。
原摘要

This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture-generation task and a text-mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large-scale user studies, collecting over 23,000 votes from 869 test-takers. In the motion-realism study, the dataset's filtered segments had substantially higher motion quality than all challenge submissions (68-95% pairwise winrate). In the speech-alignment study, the motion-capture segments provided a conceptual ceiling at 62% alignment score, with the top submission significantly behind at 32% and the rest only slightly above the 0% expected of an input-independent system. In the dyadic study, motion capture again set the ceiling at 65% appropriateness score, but no submission scored substantially above chance, indicating that the systems could not yet respond to the interlocutor. Finally, the semantic mismatching evaluation found highly expressive gestures in the dataset (test-takers identified the matching transcript 79% of the time), yet almost all submissions failed to generate semantically expressive motion, with the best achieving only an 8% appropriateness score. The collected votes and outputs will be made publicly available at https://genea-workshop.github.io/2026/challenge/ to facilitate reproducibility and further research.

Xin Dong, Vikash V. Gayah
总结: 引入大模型辅助语义站点表示以增强事件驱动公交滞站强化学习控制。
方法: 离线用大模型将异质站点信息转为固定语义嵌入并入深度Q学习控制器。
证据: 两条线路随机仿真中语义控制器较最佳Daganzo基线降低车头时距变异。
为什么适合我: 公交滞站控制与腿式人形机器人全身运动无关。
原摘要

Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding controllers primarily rely on instantaneous operational variables or route-specific stop identifiers, which provide limited information about the functional and operational context of individual stops and constrain policy reuse across routes. This study introduces an LLM-assisted semantic stop representation for event-driven bus holding control. An LLM is used offline to transform heterogeneous stop information, including physical attributes, surrounding activity context, and historical operational characteristics, into fixed semantic embeddings that are incorporated into a deep Q-learning controller without requiring real-time LLM inference. Experiments are conducted in stochastic simulations calibrated with observed data from two bus routes. Compared with the best calibrated Daganzo baseline, the semantic controller reduces headway variability, bunching events, and passenger waiting time by 32.0%, 69.2%, and 24.0%, respectively. A route-specific stop identifier does not improve the spacing-only controller, whereas semantic stop information improves headway regularity, waiting time, and holding effort, providing a more favorable overall trade-off across control objectives. Cross-route experiments further show that zero-shot transfer provides limited immediate generalization, while warm-start fine-tuning accelerates early-stage learning and improves transferred policies; cold-start training nevertheless achieves the best final performance. These findings suggest that semantic state representations can complement conventional operational states and support adaptation-based policy reuse across related transit routes.

Yidong Wang, Yan Zhan, Ziteng Feng, ... , Wei Ye, Shikun Zhang
总结: 提出TrustRoboReward多范式奖励模型框架以解决机器人奖励建模瓶颈。
方法: 用偏好有序等渗分数编辑POISE统一四范式数据集并解决偏好分数不一致。
证据: 摘要强调解决多范式训练噪声而未给出具体下游性能数字。
为什么适合我: 具身机器人奖励模型利于强化学习操控,但非腿式全身运动重点。
原摘要

Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-specific annotations. Existing open-source VLM reward judges like RoboReward adopt simple 1--5 trajectory progress scoring, lacking pairwise preferences for RLHF, DPO and Bradley-Terry frameworks, while failing to optimize video scene understanding. Augmenting RoboReward with pairwise comparison and video-QA supervision causes inconsistency between pairwise preferences and pointwise scores, introducing training noise and hurting downstream performance---an issue aggregation methods such as TrustJudge cannot resolve. To address this, we propose TrustRoboReward, a multi-paradigm reward modeling framework equipped with Preference-Ordered Isotonic Score Editing (POISE). We construct a unified four-paradigm dataset with trajectory progress scoring (Score-A), video-QA answer quality scoring (Score-B), and their pairwise counterparts (Pair-A, Pair-B). Pairwise labels align better with human judgment than pointwise scores, inspiring us to calibrate pointwise scores to avoid score-pair reversals against pairwise preferences. POISE rectifies pointwise scores and eliminates cross-paradigm reversal conflicts unresolved by TrustJudge. Theoretically, POISE reduces score-pair reversal conflicts from 20.15% to 0%, whereas TrustJudge retains 20.46% conflicts on the same corpus. Evaluated on our benchmark, Qwen3-VL-4B trained with POISE achieves an overall reward score of 77.96%, nearly matching GPT-5-mini (78.09%, gap 0.13%) and outperforming the strongest RoboReward-4B baseline by 10.13%. It also lifts test-time score-pair consistency to 71.90%, exceeding RoboReward-4B (57.26%) and GPT-5-mini (68.09%). Integrating TrustJudge aggregation during inference boosts the overall score to 78.57%, surpassing the GPT-5-mini teacher model.