Papers for 2026-02-23

10 papers
Elisa Alboni, Pietro Noah Crestaz, Elias Fontanari, Andrea Del Prete
总结: CACTO-BIC通过偏置初始状态采样提升数据效率并利用GPU加速轨迹优化,使Actor-Critic与轨迹优化结合的方法在高维控制任务中比CACTO更快、更省样本,并能在AlienGO四足机器人等实时场景中达到接近PPO的效果。
原摘要

Trajectory Optimization (TO) and Reinforcement Learning (RL) offer complementary strengths for solving optimal control problems. TO efficiently computes locally optimal solutions but can struggle with non-convexity, while RL is more robust to non-convexity at the cost of significantly higher computational demands. CACTO (Continuous Actor-Critic with Trajectory Optimization) was introduced to combine these advantages by learning a warm-start policy that guides the TO solver towards low-cost trajectories. However, scalability remains a key limitation, as increasing system complexity significantly raises the computational cost of TO. This work introduces CACTO-BIC to address these challenges. CACTO-BIC improves data efficiency by biasing initial-state sampling leveraging a property of the value function associated with locally optimal policies; moreover, it reduces computation time by exploiting GPU acceleration. Empirical evaluations show improved sample efficiency and faster computation compared to CACTO. Comparisons with PPO demonstrate that our approach can achieve similar solutions in less time. Finally, experiments on the AlienGO quadruped robot demonstrate that CACTO-BIC can scale to high-dimensional systems and is suitable for real-time applications.

Shenghong He
总结: 本文提出基于优势的对抗 Transformer(AAT),通过多尺度因果自注意力建模扰动的时间相关性并用加权优势机制引导生成更有效的对抗样本,从而在 Atari、DeepMind Control Suite 和 Google Football 等强化学习任务中达到或超过主流攻击方法的效果。
原摘要

Extensive research demonstrates that Deep Reinforcement Learning (DRL) models are susceptible to adversarially constructed inputs (i.e., adversarial examples), which can mislead the agent to take suboptimal or unsafe actions. Recent methods improve attack effectiveness by leveraging future rewards to guide adversarial perturbation generation over sequential time steps (i.e., reward-based attacks). However, these methods are unable to capture dependencies between different time steps in the perturbation generation process, resulting in a weak temporal correlation between the current perturbation and previous perturbations.In this paper, we propose a novel method called Advantage-based Adversarial Transformer (AAT), which can generate adversarial examples with stronger temporal correlations (i.e., time-correlated adversarial examples) to improve the attack performance. AAT employs a multi-scale causal self-attention (MSCSA) mechanism to dynamically capture dependencies between historical information from different time periods and the current state, thus enhancing the correlation between the current perturbation and the previous perturbation. Moreover, AAT introduces a weighted advantage mechanism, which quantifies the effectiveness of a perturbation in a given state and guides the generation process toward high-performance adversarial examples by sampling high-advantage regions. Extensive experiments demonstrate that the performance of AAT matches or surpasses mainstream adversarial attack baselines on Atari, DeepMind Control Suite and Google football tasks.

Ali Saheb Pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan, Pablo Samuel Castro
总结: 本文提出用低成本的 Sketched Isotropic Gaussian Regularization 将深度强化学习表示塑造成各向同性高斯分布,从而在非平稳目标下提升训练稳定性、适应性和性能,并减少表示坍塌与神经元休眠。
原摘要

Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data distributions evolve over time. We show that under non-stationary targets, isotropic Gaussian embeddings are provably advantageous. In particular, they induce stable tracking of time-varying targets for linear readouts, achieve maximal entropy under a fixed variance budget, and encourage a balanced use of all representational dimensions--all of which enable agents to be more adaptive and stable. Building on this insight, we propose the use of Sketched Isotropic Gaussian Regularization for shaping representations toward an isotropic Gaussian distribution during training. We demonstrate empirically, over a variety of domains, that this simple and computationally inexpensive method improves performance under non-stationarity while reducing representation collapse, neuron dormancy, and training instability.

Erik Garcia Oyono, Jialin Lin, Dandan Zhang
总结: 本文提出一种由三类嵌磁立方模块组成的模块化磁控毫米机器人平台,可通过二维时变均匀/梯度磁场实现闭环导航、自组装、多模态运动、形态重构与抓取操作,展示了其在受限环境中执行可扩展自适应任务的潜力。
原摘要

Modular small-scale robots offer the potential for on-demand assembly and disassembly, enabling task-specific adaptation in dynamic and constrained environments. However, existing modular magnetic platforms often depend on workspace collisions for reconfiguration, employ bulky three-dimensional electromagnetic systems, and lack robust single-module control, which limits their applicability in biomedical settings. In this work, we present a modular magnetic millirobotic platform comprising three cube-shaped modules with embedded permanent magnets, each designed for a distinct functional role: a free module that supports self-assembly and reconfiguration, a fixed module that enables flip-and-walk locomotion, and a gripper module for cargo manipulation. Locomotion and reconfiguration are actuated by programmable combinations of time-varying two-dimensional uniform and gradient magnetic field inputs. Experiments demonstrate closed-loop navigation using real-time vision feedback and A* path planning, establishing robust single-module control capabilities. Beyond locomotion, the system achieves self-assembly, multimodal transformations, and disassembly at low field strengths. Chain-to-gripper transformations succeeded in 90% of trials, while chain-to-square transformations were less consistent, underscoring the role of module geometry in reconfiguration reliability. These results establish a versatile modular robotic platform capable of multimodal behavior and robust control, suggesting a promising pathway toward scalable and adaptive task execution in confined environments.

Pranay Anchuri
总结: RAmmStein将集中式AMM流动性管理建模为最优脉冲控制问题,并用包含均值回复速度等市场状态的深度强化学习学习“何时不动、何时再平衡”,在真实高频数据中以显著更少的再平衡和更低Gas成本取得优于其他现实策略的净收益。
原摘要

Concentrated liquidity provision in decentralized exchanges presents a fundamental Impulse Control problem. Liquidity Providers (LPs) face a non-trivial trade-off between maximizing fee accrual through tight price-range concentration and minimizing the friction costs of rebalancing, including gas fees and swap slippage. Existing methods typically employ heuristic or threshold strategies that fail to account for market dynamics. This paper formulates liquidity management as an optimal control problem and derives the corresponding Hamilton-Jacobi-Bellman quasi-variational inequality (HJB-QVI). We present an approximate solution RAmmStein, a Deep Reinforcement Learning method that incorporates the mean-reversion speed (theta) of an Ornstein-Uhlenbeck process among other features as input to the model. We demonstrate that the agent learns to separate the state space into regions of action and inaction. We further extend the framework with RAmmStein-Width, which jointly optimizes rebalancing timing and position width via a 6-action DDQN. We evaluate the framework using high-frequency 1Hz Coinbase trade data comprising over 6.8M trades on a realistic environment (10M TVL, 1% default width). Experimental results show that RAmmStein achieves a net ROI of 1.60%, the highest among all realistic (non-omniscient) strategies, while greedy strategies lose up to -8.4% to gas costs. Notably, the agent reduces rebalancing frequency by 85% compared to greedy rebalancing. RAmmStein-Width discovers extreme parsimony on its own, executing only 9 rebalances and $40 in gas, and degrades more slowly than all active strategies at elevated gas costs. Our results demonstrate that regime-aware laziness can significantly improve capital efficiency by preserving the returns that would otherwise be eroded by the operational costs.

Hong Wang, Xuwei Fan, Zhipeng Cheng, ... , Minghui Liwang, Xiaoyu Xia
总结: 本文提出一种面向隐私感知边缘设备协同DNN推理的安全多智能体深度强化学习框架HC-MAPPO-L,通过分层策略与拉格朗日约束优化模型部署、用户关联、模型切分和资源分配,在满足长期时延约束的同时降低能耗与隐私成本。
原摘要

As Deep Neural Network (DNN) inference becomes increasingly prevalent on edge and mobile platforms, critical challenges emerge in privacy protection, resource constraints, and dynamic model deployment. This paper proposes a privacy-aware collaborative inference framework, in which adaptive model partitioning is performed across edge devices and servers. To jointly optimize inference delay, energy consumption, and privacy cost under dynamic service demands and resource constraints, we formulate the joint problem as a Constrained Markov Decision Process (CMDP) that integrates model deployment, user-server association, model partitioning, and resource allocation. We propose a Hierarchical Constrained Multi-Agent Proximal Policy Optimization with Lagrangian relaxation (HC-MAPPO-L) algorithm, a safe reinforcement learning-based framework that enhances Multi-Agent Proximal Policy Optimization (MAPPO) with adaptive Lagrangian dual updates to enforce long-term delay constraints. To ensure tractability while maintaining coordination, we decompose the CMDP into three hierarchically structured policy layers: an auto-regressive based model deployment policy, a Lagrangian-enhanced user association and model partitioning policy, and an attention-based resource allocation policy. Extensive experimental results demonstrate that HC-MAPPO-L consistently satisfies stringent delay constraints while achieving a superior balance among energy consumption and privacy cost, outperforming representative baseline algorithms across varying problem scales and resource configurations.

Tamil Selvan Gurunathan, Aryya Gangopadhyay
总结: 该论文提出将希尔伯特空间填充曲线先验融入DQN/PPO的去中心化多机器人覆盖与探索框架,以减少稀疏奖励环境中的冗余并提升覆盖效率、收敛速度和可扩展性,同时通过可执行的SE(2)轨迹接口在Spot机器人上验证了其实际可行性。
原摘要

We present a coverage framework that integrates Hilbert space-filling priors into decentralized multi-robot learning and execution. We augment DQN and PPO with Hilbert-based spatial indices to structure exploration and reduce redundancy in sparse-reward environments, and we evaluate scalability in multi-robot grid coverage. We further describe a waypoint interface that converts Hilbert orderings into curvature-bounded, time-parameterized SE(2) trajectories (planar (x, y, θ)), enabling onboard feasibility on resource-constrained robots. Experiments show improvements in coverage efficiency, redundancy, and convergence speed over DQN/PPO baselines. In addition, we validate the approach on a Boston Dynamics Spot legged robot, executing the generated trajectories in indoor environments and observing reliable coverage with low redundancy. These results indicate that geometric priors improve autonomy and scalability for swarm and legged robotics.

Yarden As, Dhruva Tirumala, René Zurbrügg, ... , Andreas Krause, Markus Wulfmeier
总结: 该论文通过在三种真实机器人平台上进行100次在线强化学习训练,系统分析算法、系统与实验设计选择的影响,指出常见默认设置可能有害,并总结出一组更稳健、易采用的实践以降低真实机器人在线RL部署难度。
原摘要

We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots. Across 100 real-world training runs on three distinct robotic platforms, we systematically ablate algorithmic, systems, and experimental decisions that are typically left implicit in prior work. We find that some widely used defaults can be harmful, while a set of robust, readily adopted design choices within standard RL practice yield stable learning across tasks and hardware. These results provide the first large-sample empirical study of such design choices, enabling practitioners to deploy online RL with lower engineering effort.

Nathan Gavenski, Felipe Meneguzzi, Odinaldo Rodrigues
总结: 本文主张模仿学习应从追求精确复现转向终身可组合适应性,通过学习可重组的行为基元、建立泛化指标并设计混合架构,使智能体能在新情境和目标变化中无需重新训练地灵活行动。
原摘要

Imitation learning stands at a crossroads: despite decades of progress, current imitation learning agents remain sophisticated memorisation machines, excelling at replay but failing when contexts shift or goals evolve. This paper argues that this failure is not technical but foundational: imitation learning has been optimised for the wrong objective. We propose a research agenda that redefines success from perfect replay to compositional adaptability. Such adaptability hinges on learning behavioural primitives once and recombining them through novel contexts without retraining. We establish metrics for compositional generalisation, propose hybrid architectures, and outline interdisciplinary research directions drawing on cognitive science and cultural evolution. Agents that embed adaptability at the core of imitation learning thus have an essential capability for operating in an open-ended world.

Shirui Chen, Cole Harrison, Ying-Chun Lee, ... , Dieter Fox, Ranjay Krishna
总结: TOPReward通过利用预训练视频视觉语言模型的内部token概率零样本估计机器人任务进展,在130多个真实任务中显著优于现有方法,并可用于成功检测和奖励对齐的行为克隆。
原摘要

While Vision-Language-Action (VLA) models have seen rapid progress in pretraining, their advancement in Reinforcement Learning (RL) remains hampered by low sample efficiency and sparse rewards in real-world settings. Developing generalizable process reward models is essential for providing the fine-grained feedback necessary to bridge this gap, yet existing temporal value functions often fail to generalize beyond their training domains. We introduce TOPReward, a novel, probabilistically grounded temporal value function that leverages the latent world knowledge of pretrained video Vision-Language Models (VLMs) to estimate robotic task progress. Unlike prior methods that prompt VLMs to directly output progress values, which are prone to numerical misrepresentation, TOPReward extracts task progress directly from the VLM's internal token logits. In zero-shot evaluations across 130+ distinct real-world tasks and multiple robot platforms (e.g., Franka, YAM, SO-100/101), TOPReward achieves 0.947 mean Value-Order Correlation (VOC) on Qwen3-VL, dramatically outperforming the state-of-the-art GVL baseline which achieves near-zero correlation on the same open-source model. We further demonstrate that TOPReward serves as a versatile tool for downstream applications, including success detection and reward-aligned behavior cloning.