Papers for 2026-09-06

10 papers

Latent Energy Action Planning with World Models

2.0/5 偏低 裁判分 7.0 Strong 当日相对
Phu Pham, Aniket Bera
总结: LEAP用潜在能量规划动作以匹配目标潜在与描述符。
方法: 动作作可微变量,经冻结世界模型优化并用准牛顿求解。
证据: 四控制域用官方LeWM检查点提升平均性能。
为什么适合我: 潜在世界模型MPC可支持复杂地形全身运动规划。
推荐理由: 潜空间世界模型上的可微动作规划与终端目标匹配,方法上靠近模型预测控制,但对象是通用控制域,不是人形全身跟踪、腿足跑酷或loco-manipulation。
原摘要

Latent world models support efficient model predictive control from high-dimensional observations, yet optimizing a single learned latent objective can favor action sequences whose decoder-predicted terminal descriptor does not match the goal descriptor. We introduce Latent Energy Action Planning (LEAP), which treats the complete action horizon as a differentiable variable and optimizes it through a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal-window state energy. Low energy requires the predicted terminal latent to agree with the goal latent and the decoder-predicted terminal descriptor to agree with the goal descriptor. A frozen goal-conditioned proposal initializes the search, a quasi-Newton solver refines actions through the autoregressive rollout, and post-optimization projection enforces the admissible action range. Across four control domains using the officially released LeWM checkpoints, the complete LEAP planning system raises mean success from 77.5% for LeWM planned with the cross-entropy method (LeWM+CEM) to 94.8% under a matched protocol, a 17.3-percentage-point improvement, while retaining the frozen LeWM representation and predictor.

Amal Dev Haridevan, Junjie Kang, Jinjun Shan
总结: GzDRL提供无中间件确定性Gazebo强化学习框架。
方法: 直接同步动作与物理更新,实现可复现高吞吐训练。
证据: 工作站吞吐最高,支持四足真机策略迁移。
为什么适合我: 可复现RL利于腿足机器人训练与真机部署。
推荐理由: Gazebo上可复现的深度强化学习仿真框架,只与机器人RL基础设施沾边,不涉及全身运动、接触丰富控制或sim-to-real运动策略。
原摘要

We present GzDRL, a novel single-process reinforcement learning (RL) framework for Gazebo that overcomes longstanding bottlenecks in scalable, reproducible robotics experimentation. Unlike conventional middleware-based RL-Gazebo integrations that suffer from nondeterminism and irreproducibility, GzDRL introduces a systematic, middleware-free environment-stepping mechanism that directly synchronizes agent actions and physics updates. This design enables deterministic, high-throughput data collection, efficient vectorization, and reproducible RL training and evaluation. Comprehensive benchmarks demonstrate that GzDRL achieves the highest workstation throughput among the evaluated frameworks while remaining competitive with GPU-accelerated simulators on laptop hardware, and maintains precise agent-environment synchronization, multi-agent scalability, and experiment-level reproducibility. We further validate sim-to-real transfer by deploying learned policies directly onto a physical quadrotor, without fine-tuning. Our results establish GzDRL as an accessible and reproducible platform for advancing RL in robotics and automation.

Haoyu Wang, Songchun Zhang, Haoran Li, ... , Zeyue Xue, Nan Duan
总结: 基于虚幻引擎构建动作条件多视角视频数据管道。
方法: 两阶段:实时物理记录轨迹,离线高质量渲染。
证据: 分布式系统支持大规模动作对齐视频生成。
为什么适合我: 动作条件数据可预训练世界模型支持运动生成。
推荐理由: 用虚幻引擎生成带控制信号的角色多视角视频,靠近物理角色数据与动作条件世界模型,但不是可跟踪、可重定向或可上真机的全身控制。
原摘要

Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is difficult to obtain from ordinary real-world video because the actions that caused each visual change are typically unknown. We present a large-scale synthetic data production pipeline built on Unreal Engine for generating action-conditioned, multi-view video. To accommodate the different execution requirements of real-time physics and high-quality offline rendering, the pipeline executes trajectory generation and final rendering in two stages: Stage I runs real physics in PIE and records per-frame character states, control inputs, and camera states into an intermediate trajectory representation; Stage II replays those trajectories in a new engine process and renders them offline with Movie Render Queue (MRQ). Around this core, we develop a distributed production system with cache-aware task partitioning, node-local slot scheduling, automated scene screening, aesthetic and luminance filtering, partial-output recovery, asynchronous upload, and continuous cluster health monitoring. The production cluster contains 25 servers with eight NVIDIA RTX 5090 GPUs per server. From 2,384 asset packs, 429 levels were retained for production together with a pool of 40 humanoid characters. The pipeline has produced 2,691 hours of 1080p video and 6,076 hours of 720p video. We describe the system architecture, the implementation decisions that emerged from production failures, and the limitations of using perceptual quality proxies for world-model data curation. The pipeline described in this report constitutes the Unreal Engine synthetic-data production component used in EchoWM.

Tam W. Nguyen
总结: 泰勒知情间接自适应预测控制用雅可比冻结仿射预测器。
方法: 有限泰勒展开近似,RLS在线辨识,冻结雅可比做MPC。
证据: 不稳定非线性基准仿真显示高阶提升跟踪精度。
为什么适合我: 自适应MPC可应用于全身模型预测控制。
推荐理由: 基于泰勒展开与冻结雅可比的间接自适应预测控制,方法上沾MPC,但是通用非线性采样系统数值例子,不是腿足或人形全身MPC。
原摘要

This paper develops a Taylor-informed indirect adaptive predictive control framework for nonlinear sampled-data systems using Jacobian-frozen affine predictors. A finite Taylor expansion approximates the sampled nonlinear dynamics, and recursive least squares (RLS) identifies its polynomial coefficients online. At each sampling instant, the Jacobian of the identified map is evaluated at the current operating point and frozen over the prediction horizon, yielding an affine predictor for model predictive control. In contrast to generic nonlinear feature dictionaries, the implemented polynomial dictionary is a forward-Euler/Taylor-structure-informed reduced dictionary. Exact joint-odd symmetry eliminates even-total-degree monomials, whereas additional forward-Euler-informed pruning constitutes a deliberate model reduction. Numerical simulations on an unstable nonlinear benchmark compare different Taylor degrees. The results show that higher-order models improve tracking accuracy as the operating point moves farther from the expansion point while maintaining comparable control effort. The complete MATLAB implementation is publicly available to facilitate reproducibility.

Is One Step Enough for Offline Policy Improvement?

2.0/5 偏低 裁判分 2.0 Candidate 当日相对
Soohyun Choi, Seonvin Cho, Songnam Hong
总结: 研究多步近端策略改进如何组成离线策略提升。
方法: 参数化总视界与阶段,重中心化近端目标进行改进。
证据: TD3+BC在D4RL运动上显示细分拓宽有用范围。
为什么适合我: 离线策略改进可增强机器人模仿与强化学习。
推荐理由: 离线强化学习中多步近端策略改进的理论分析,方法泛相关,但没有机器人运动、模仿或真机对象。
原摘要

Behavior regularization in offline reinforcement learning limits the exploitation of critic errors, but strong anchoring can also restrict policy improvement. We study how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy. We parameterize the procedure by a nominal total horizon $T$ and $K$ stages with local horizon $T/K$, distinguishing subdivision at a fixed total horizon from additional refinement at a common local horizon. Our analysis shows that sequential re-centering can reach endpoints unavailable to any single proximal step and characterizes how subdivision reduces the leading local discretization error of ideal updates under a fixed critic. We consider TD3+BC and IQL-based policy extraction to examine how improvement composition interacts with actor objectives and policy geometry. TD3+BC experiments on D4RL locomotion suggest that subdivision can broaden the range of useful total horizons, while adding refinement stages at a fixed small local horizon can improve return. The results identify improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and additional policy extraction.

Symmetries and Causality: Causal Effect Identification Beyond IID Data

1.0/5 偏低 裁判分 5.0 Candidate 当日相对
Martin Rabel, Jakob Runge
总结: 基于对称性提出超越IID数据的因果效应识别语言。
方法: 用数据对称性保持因果机制不变,形式化模型与查询。
证据: 匹配标准IID结果并扩展复杂因果查询范围。
为什么适合我: 因果推理可助力世界建模与复杂运动智能。
推荐理由: 对称性与因果识别的抽象理论,仅把强化学习世界模型当作动机,与人形/腿足运动智能无关。
原摘要

In the natural sciences, symmetries and cause-effect relationships are ubiquitous. Yet for complex machine-learning tasks, like world-modeling in reinforcement learning, they appear difficult to harness. We propose a formal description of statistical systems based on symmetries in data leaving causal mechanisms invariant. The result is an abstract, simple and general mathematical language for causal reasoning. This paper provides formal descriptions of models and queries, setting up this language, and the formal infrastructure and strategies for their mathematically rigorous identification from data within this formalism. This approach reproduces and matches standard theoretical results on IID data and transport of experimental and non-experimental data. But its main purpose is to unify and substantially extend the scope of causal reasoning, in going beyond IID data and in approaching complex causal queries not captured by do- or soft-interventions. This new perspective on causally relevant aspects of data-modeling additionally sheds new light on well-known structures like c-components or hedges but also includes aspects of missing data and is inherently well-suited for the description of transfer and robustness properties.

The Dually Flat Geometry of Planning as Inference

1.0/5 偏低 裁判分 4.0 Candidate 当日相对
Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf
总结: 规划即推理的对偶平坦几何用访问测度刻画。
方法: 嵌入重置规划过程,访问测度形成对偶平坦流形。
证据: 自然梯度步求解非线性泛函,TD误差为边际效用。
为什么适合我: 几何视角可深化全身控制与强化学习理论。
推荐理由: 规划即推断的对偶平坦信息几何,属于强化学习与理论神经科学,不对应全身控制或运动先验。
原摘要

We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

1.0/5 偏低 裁判分 3.0 Candidate 当日相对
Sijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi
总结: LeanGRPO消除扩散强化学习中冗余重计算。
方法: 重构数据并行布局,引入无重计算训练调度。
证据: 保留与重加权调度避免重算并降低开销。
为什么适合我: 高效扩散RL可用于人体动作生成与重定向。
推荐理由: 面向图像和视频生成模型的扩散强化学习重计算消除,不是机器人运动扩散或模仿学习。
原摘要

Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO, recompute selected timesteps with gradient tracking after rollout. Under on-policy training with the same backend for rollout and update, this recomputation is mathematically redundant. Intuitively, the rollout and policy update steps can reuse the same feed-forward backbone to avoid redundant computation, but doing so can incur a large memory overhead during rollout. To address the issue, we present LeanGRPO by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL: (1) LeanGRPO-Retain enables gradient tracking during rollout and directly reuses the resulting computation graphs and saved activations for backward during update, requiring no recomputation; and (2) LeanGRPO-Reweight also enables gradients during rollout, but immediately backpropagates each selected step using a provisional advantage and delays gradient synchronization, then corrects the provisional gradients with the true advantage after the trajectory is completed. These schedules target different model scales and input sizes. Across FlowGRPO/DanceGRPO with FLUX.1-dev and Wan, LeanGRPO achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.

Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk
总结: DRACO用动态量规细粒度信用分配训练长时程智能体。
方法: 动态生成量规,轨迹评分后闭式重分配到步骤优势。
证据: AppWorld上比基模提升15.9点,比稀疏奖励GRPO高5.3点。
为什么适合我: 长时程信用分配利于loco-manipulation任务训练。
推荐理由: 长程智能体的细粒度信用分配,评测在AppWorld,与机器人运动控制无关。
原摘要

Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO. The redistribution is closed-form and does not introduce any trained attribution module. On AppWorld, DRACO gains 15.9 points over the base model and 5.3 points over GRPO trained with a sparse ground-truth reward, despite not using any verifiers itself. On out-of-domain Tau-Bench, it gains 5.3 points over the base model even without a frontier judge, beating both ground-truth-reward training and other rubric-based training settings. The code for DRACO is available at https://github.com/IBM/draco.

Junlong Wu, Jiuzhou Lin, Jia Sun, ... , Fan Yang, Tingting Gao
总结: RA-GRPO通过反射感知偏好优化改进视觉生成。
方法: 扩散反射反转过程纠正轨迹,引导潜在到高概率区。
证据: 改善语义忠实与视觉真实,避免局部最优。
为什么适合我: 反射优化可生成可跟踪可重定向的运动轨迹。
推荐理由: 视觉生成的反思式偏好优化,属于文生图/视频对齐,不在兴趣范围内。
原摘要

Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based methods often explore inefficiently, making them vulnerable to local optima that may degrade semantic faithfulness and visual realism. To address these challenges, we present Reflection-Aware GRPO (RA-GRPO), a new RL-based preference alignment framework for diffusion generative models. The core idea is to improve "forward" generation by incorporating "backward" reflection during optimization. We first introduce Diffusion Reflection, which rectifies intermediate sampling trajectories by inverting the diffusion process with a weak estimator, guiding latent states toward higher-probability regions of the true data manifold. Furthermore, we introduce Counterfactual Path Synthesis to implicitly distill these rectified trajectories into the policy, enabling the model to internalize the benefits of search-based exploration without incurring inference-time overhead. Extensive experiments on T2I and T2V models demonstrate that RA-GRPO significantly outperforms existing methods, particularly in mitigating reward hacking and improving generalization. The method remains architecture-agnostic and integrates seamlessly with standard pipelines, suggesting a promising direction for stable preference alignment.