Papers for 2026-07-26

10 papers

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

2.0/5 偏低 裁判分 2.0 Candidate 当日相对
Mengfei Zhao, Dihong Huang, Yikai Tang, ... , Jianfei Yang, Jiachen Li
总结: 社区驱动可扩展机器人操作数据引擎与基准。
方法: 浏览器遥操作收集演示并自动生成验证增强数据。
证据: 含207任务超5万轨迹,统一评估VLA策略缩放。
为什么适合我: 可扩展操作数据支持人形loco-manipulation模仿学习。
推荐理由: 通用操作演示采集与数据引擎,仅弱相关于遥操作数据采集,并非人形全身或腿足loco-manipulation。
原摘要

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalable robot learning, which enables browser-based teleoperation for large-scale demonstration collection, automatically generates and validates new manipulation tasks, and transforms community-collected demonstrations into training-ready data through automated success checking, quality filtering, trajectory smoothing, and visual and physics-based augmentation. The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories. Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-out protocol. We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyze scaling behavior across different data volumes. Continual pretraining on AXIS substantially improves the overall success rate of $π_{0.5}$ by 5.8%, outperforms the model pretrained on RoboCasa365 by 37.3%, and exhibits consistent scaling with increasing data volume, with the largest gains observed under layout, sensor-noise, and camera perturbations.

World models of environment, agent and joint agent-environment systems

1.0/5 偏低 裁判分 3.0 Candidate 当日相对
Manuel Baltieri, Filippo Torresan, Yivan Zhang, Alexander Boyd, Fernando E. Rosas
总结: 区分环境智能体与联合系统的世界模型通道。
方法: 用计算力学定义ε-转换器作为规范预测模型。
证据: 规范模型恢复预测状态表示并扩展至智能体联合。
为什么适合我: 世界模型助力复杂地形感知运动与全身MPC。
推荐理由: 计算力学视角的世界模型与ε-machine理论,不涉及人形/腿足全身控制、运动跟踪或loco-manipulation。
原摘要

World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they predict, such as observations, rewards, states, latent or information states. We argue that there is a prior distinction: which channel they model. We consider three cases: the environment channel $O_{:} \mid A_{:}$, the agent channel $A_{:} \mid O_{:}$, and the realised joint process $(A, O)_{:}$, equivalently viewed as a channel with no inputs. Using computational mechanics, we define canonical predictive models for these three cases as $ε$-transducers or $ε$-machines. Canonical environment models recover standard predictive state representations, while the other two give analogous notions of canonical models for the agent and the joint system. We then build canonical support-restricted environment and agent models induced by closed-loop coupling, whose predictive equivalences range over continuations supported by the realised interaction. The key structural result is that canonical support-restricted environment states factor through the canonical joint causal states, and their transition structure is induced directly from the joint model; the agent-side construction is dual. Finally, we give a POMDP/controller example in which the unrestricted environment model has infinitely many states while the canonical support-restricted model induced by the coupling is finite. The framework clarifies what different world models are models of, and how coupling and support restriction can change their canonical predictive structure and complexity.

Adaptive Multi-Horizon Reinforcement Learning

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Manoosh Samiei, Doina Precup, Paul Masset
总结: 自适应多时域强化学习平衡短长期后果。
方法: 自适应选择组合多折扣时域无需手动调参。
证据: MiniGrid及持续三任务切换中识别有效折扣因子。
为什么适合我: 多时域适应利于跑酷接触丰富场景持续学习。
推荐理由: 多时域折扣的通用强化学习,实验在MiniGrid,与全身运动智能无关。
原摘要

Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents exhibit flexible and adaptive temporal discounting, suggesting that effective planning requires multiple timescales. Here, we propose a multi-horizon approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changes in reward structure without manual discount-factor tuning. This flexibility makes the method particularly suitable for continual learning scenarios involving task switches and varying environmental configurations. Empirically, we demonstrate that our approach identifies effective discount factors across a range of MiniGrid environments, including continual settings composed of three sequentially changing tasks. These results suggest that adaptive temporal discounting can improve parameter efficiency and enhance adaptability in both artificial and biologically inspired learning systems.

Robust Asynchronous Q-Learning under Reward and State Corruption via Batching

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Sreejeet Maity, Aritra Mitra
总结: 奖励状态腐败下鲁棒异步Q学习算法。
方法: 分批数据构建鲁棒贝尔曼算子抗Huber污染。
证据: 高概率误差界匹配标准Q学习加小腐败项。
为什么适合我: 真机噪声鲁棒性利于腿足全身控制部署。
推荐理由: 抗奖励与状态污染的异步Q-learning理论,无机器人运动或接触控制内容。
原摘要

Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both the reward and state observations of the learner following the Huber contamination model. To defend against such data corruption, we propose BR-Async-Q: a novel, epoch-based, robust Q-learning algorithm built upon two key ideas: (i) partitioning the online data stream into batches to reduce variance, and (ii) constructing robust estimates of the Bellman optimality operator using such batched data. We prove a high-probability $\ell_\infty$ error bound for BR-Async-Q that matches that for vanilla Q-learning, up to a small additive term that scales with the fraction of corrupted samples. To our knowledge, this provides the first robustness guarantee for asynchronous Q-learning subject to both reward and state corruption. Furthermore, when only rewards are corrupted, the dependence of our algorithm's bound on the corruption fraction is minimax optimal.

Hongju Pae
总结: 主动推理中视角潜变量促成因果涌现。
方法: 分离快感知潜z与慢全局潜g由预测误差驱动。
证据: Φr集中于g,学习使解耦变正且体制不变。
为什么适合我: 主动推理启发人形感知运动分层控制架构。
推荐理由: 主动推断与因果涌现的信息论分析,不对应感知运动、模仿或全身MPC。
原摘要

A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that $Φ_r$ grows with training and tracks reward improvement. For active inference, this raises the question of how reward-free predictive organization relates to such information-theoretic signatures. I test this within an active inference agent whose architecture separates a fast perception latent $z$ from a slow global latent $g$, where $g$ is driven by prediction error and structurally decoupled from policy gradients. In a reward-free environmental regime-switching protocol, $Φ_r$ concentrates in $g$; its aggregate magnitude is largely architectural and decreases with training. The substantive effect of learning becomes legible only at the atom-compositional level: decoupling flips sign from negative to positive and becomes regime-invariant under environmental change, while downward causation carries the regime-dependent adjustment. These results identify $g$ as the architectural locus of $Φ_r$-relevant temporal organization in an active inference agent, and argue against reading scalar $Φ_r$ as a direct index of learned integration.

Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Dzmitry Malyshau
总结: 冻结视觉特征紧凑行为克隆玩Quake游戏。
方法: 六层变压器在冻结DINOv3编码器上行为克隆。
证据: 每集达关键区域,19/20有击杀未通关优于基线。
为什么适合我: 简单模仿学习可借鉴人体动作全身控制迁移。
推荐理由: Quake第一人称游戏的行为克隆,不是物理角色或人形运动跟踪。
原摘要

We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a six-layer transformer over a frozen DINOv3 encoder. It is trained on the Quake subset of the public Pixels2Play corpus: 6,849 recordings (about 474.7 hours), represented as 17.09 million cached decision frames with keyboard and mouse actions. One sampled training epoch uses 517,048 four-frame windows and takes 3.3 minutes of policy-head optimization on one RTX 5080, excluding one-time feature extraction. We evaluate two independent batches of 20 stochastic, 120-second episodes on Quake E1M1. Cortex does not complete the level, but every episode reaches the opening door, button room, and gate descent; 19 of 20 episodes in each batch record at least one kill. Under the same time-controlled harness, released P2P-150M and NitroGen checkpoints remain shallower in five matched-duration episodes each. These comparisons are limited by small reference samples and different native interfaces. Ablations show that denser visual tokens improve combat and survival, while longer optimization and naive action history improve offline metrics without consistently improving play. The remaining failures are consistent with covariate shift and motivate targeted corrective data. We release the policy implementation, checkpoint, and a representative rollout.

Relative Value Learning

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Marc Höftmann, Jan Robine, Stefan Harmeling
总结: 直接学习相对价值差而非绝对状态价值。
方法: 反对称Δ与成对贝尔曼算子重构无偏R-GAE。
证据: 结合PPO在Atari 49游戏达竞争性能。
为什么适合我: 相对价值利于强化学习全身运动策略优化。
推荐理由: 相对价值学习与Atari上的PPO,属于通用RL,不涉及腿足或全身控制。
原摘要

In reinforcement learning, critics typically estimate absolute state values $V(s)$, estimating how good a particular situation is in isolation. However, it turns out that only differences in value are relevant for control. Motivated by this, we propose Relative Value Learning (RV), a framework that learns value differences directly via an antisymmetric function $Δ(s_i, s_j) = V(s_i) - V(s_j)$. We introduce a pairwise Bellman operator and prove it is a $γ$-contraction with a unique fixed point equal to the true value differences, derive well-posed $1$-step, $n$-step and $λ$-return targets and reconstruct generalized advantage estimation from pairwise differences to obtain an unbiased policy-gradient estimator (R-GAE). Beyond theoretical results, we integrate RV with PPO and achieve competitive performance on the Atari benchmark (49 ALE games) compared to standard PPO, indicating that relative value estimation is an effective alternative to absolute critics.

Jian Hu, Huiying Li, Hao Zhang, ... , Jan Kautz, Yi Dong
总结: 可扩展PyTorch原生智能体强化学习框架。
方法: 异步重叠rollout优化支持万亿参数模型。
证据: 验证万亿参数完整RL管道与30B持续学习。
为什么适合我: 框架加速人形机器人大规模强化学习实验。
推荐理由: 面向LLM智能体的分布式训练框架,与机器人运动学习无关。
原摘要

Agentic reinforcement learning requires rapid experimentation with agents and learning algorithms, yet large policies and long, multimodal trajectories demand substantial distributed infrastructure. We present MOLT, a lightweight, PyTorch- and Hugging Face-native framework that brings these goals together through four contributions. MOLT combines direct loading of Hugging Face models with experimentally validated trillion-parameter scalability in approximately 9.2K lines of framework code. Unified OpenAI- and Anthropic-compatible interfaces integrate existing agents with automatic handling of context compaction. Fully asynchronous training overlaps agent rollouts and policy optimization, accommodating variable agent execution times. Distributed experience storage removes centralized rollout-memory bottlenecks for long, multimodal trajectories. We experimentally validate the complete RL training pipeline on a one-trillion-parameter policy and demonstrate sustained learning with a 30B mixture-of-experts agent, establishing MOLT as a lightweight foundation for large-scale agentic RL research.

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Yipeng Shi, Zhipeng Ma, Yue Wang, ... , Peng Chen, Zhengzhou Zhu
总结: 策略感知训练脚手架提升智能体强化学习。
方法: rollout转证据卡动态调整上下文指导弱策略。
证据: 弱策略获指导,强后移除冗余保留有用变异。
为什么适合我: 自适应脚手架助复杂loco-manipulation探索。
推荐理由: 长程LLM智能体的训练脚手架,不落在模仿学习或全身控制方向。
原摘要

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by optimizing, filtering, or internalizing reusable skills. However, they remain centered on the skills themselves rather than being designed as adaptive training-time support for the evolving policy. To address this, we propose a policy-centric training paradigm that reframes skills as a dynamic training scaffold. Our framework, PATS, converts rollout groups from the latest policy into evidence cards and uses task-specific evaluation to adjust the context used in subsequent rollouts. Concrete guidance helps weak policies to complete challenging tasks. As policy improves, redundant context is revised or removed to reduce reliance on explicit guidance while preserving useful rollout variation. The policy is optimized with environmental rewards using standard RLVR, and the training scaffold is discarded at deployment. Across ALFWorld, WebShop, and seven search-augmented QA benchmarks, PATS achieves performance competitive with SOTA baselines while using 25%-50% fewer tokens.

Liu Zai, Yumeng Wang, Junchen Fu, Joemon M. Jose
总结: 荣格认知功能激活转向控制LLM人格。
方法: 荣格协议与2100+叙述数据集提取转向向量。
证据: Llama上单调控制八功能,中层集中人格信息。
为什么适合我: 转向几何或启发运动先验与动作重定向控制。
推荐理由: 用激活转向控制LLM人格,与运动先验和角色控制无关。
原摘要

Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing character narrations. Activation steering vector extraction and evaluation experiments on Llama-3.1-8B demonstrate effective monotonic control over all eight cognitive functions through activation steering. Beyond controllability, our analysis reveals that: 1. personality information is concentrated in middle transformer layers; 2. steering vectors exhibit structured geometric relationships consistent with distinctions between rational and irrational functions; 3. effective multi-dimensional steering directions cannot be recovered as linear combinations of single-function directions. These findings provide new insights into the representation of personality in LLM activation space and establish a framework for studying interpretable, effective, and multi-dimensional personality control.