Papers for 2026-08-11

10 papers
Yilei Hua, Beibei Jing, Ce Zheng, ... , Yawei Luo, Wei Yang
总结: 指令驱动3D人体运动编辑根植于文本生成过程。
方法: 闭环合成验证建Omni-MoEdit数据集,统一潜在流匹配模型,SAFE推理。
证据: 克服适应控制不佳与三元组数据规模语义有限瓶颈。
为什么适合我: 人体动作编辑支持可跟踪可重定向可上真机的全身控制。
推荐理由: 最接近人体动作生成与运动先验:指令驱动的三维人体动作编辑和flow matching,但未涉及可跟踪、重定向或上真机的全身控制。
原摘要

Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.

Nakgyu Yang, KwangBin Lee, SooJean Han
总结: 拓扑图引导扩散规划在结构层强制安全。
方法: 高层拓扑图规划器引导低层扩散模型生成连续轨迹。
证据: 理论证流形破裂条件,实验提升无碰撞目标到达率。
为什么适合我: 安全扩散规划适用复杂地形接触丰富场景运动智能。
推荐理由: 仅在扩散规划与安全引导上沾边,对象是通用轨迹流形而非人形/腿足全身运动或接触丰富控制。
原摘要

Many diffusion-based planners enforce safety through inference-time guidance, but such interleaved trajectory deformations often degrade kinematic feasibility due to manifold rupture. We propose Graph-Guided Safe Diffuser (G2SD), a hierarchical framework that leverages a high-level topological graph planner to guide a low-level diffusion model. G2SD enforces safety at a structural level by abstracting the data manifold into a learned latent graph, on which high-level planning is performed. Continuous trajectories are generated by diffusion planners, which are conditioned on the graph node representations selected by the high-level planner. Theoretical analyses demonstrate conditions under which manifold rupture occurs in diffusion planners, and show that G2SD improves safety by reducing the constraint violation probability as the number of segments increases. Experiments demonstrate that G2SD substantially outperforms baselines, increasing goal-reaching rate without any collision from 40-50% to 98% in Maze2D navigation and also achieving superior task scores in locomotion.

Donghu Kim, Youngdo Lee, Hojoon Lee, ... , Jaegul Choo, Clare Lyle
总结: 视觉连续控制中架构设计显著提升RL样本效率。
方法: 基于SAC加数据增强,加归一化层与点卷积稳定简化。
证据: 尽管简单,匹配或超越现有视觉RL方法效率。
为什么适合我: 视觉RL架构可增强人形机器人感知运动学习效率。
推荐理由: 是视觉连续控制的网络结构改进,并非腿足高程图感知、跑酷或全身loco-manipulation。
原摘要

Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challenge is pronounced in visual RL, where high-dimensional inputs often obscure learning signals. While prior work in visual RL has focused on algorithmic solutions, such as better dynamics models or exploration strategies, recent advances in state-based RL show that architectural design alone can lead to significant gains in sample efficiency. This raises an important question: Can these architectural principles transfer to visual RL? In response, we introduce V-Simba, a simple yet effective visual RL architecture inspired by the Simba architecture from state-based RL. Built on top of Soft Actor-Critic (SAC) with data augmentation, V-Simba modifies the architecture by adding normalization layers to stabilize training and using pointwise convolutions to reduce computation. Despite its simplicity, V-Simba matches or outperforms the state-of-the-art methods across the DMC, Adroit, and Meta-World benchmarks, while being more computationally efficient than DrQ-v2. We make our code publicly available at https://github.com/DAVIAN-Robotics/V-Simba.

Shiting Gong, Jianpeng Yao, Jinfeng Wang, Marco Pavone, Jiachen Li
总结: 人群跟随中分解接近与安全冲突为多约束。
方法: 稀疏任务奖励加独立成本约束,阈值显式调控权衡。
证据: 量化预测不确定性,实现可调接近安全平衡。
为什么适合我: 约束分解RL利于接触丰富loco-manipulation安全控制。
推荐理由: 人群中跟人导航的约束强化学习,属于移动社交导航,而非腿足感知运动、跑酷或全身操作。
原摘要

Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedestrians and obstacles. This conflict becomes more severe in dense scenarios, where aggressive following risks collisions and conservative margins lead to target loss, especially when pedestrian behaviors are unfamiliar or unpredictable. Existing reinforcement learning (RL) methods typically encode these competing objectives into a single dense reward, but the resulting proximity-safety balance is implicit and difficult to adjust across conditions. To address this, we decompose the human-following task into a sparse task reward and independent cost constraints within a multi-constraint RL formulation, where each constraint is managed through cost thresholds with direct behavioral meaning rather than implicit reward weight ratios, allowing explicit and tunable control over the trade-off. We further quantify the prediction uncertainty of human motions and integrate these estimates into the RL costs to enhance safety under unpredictable conditions. Extensive experiments across both in-distribution and out-of-distribution settings demonstrate that our method achieves an effective proximity-safety balance compared to baselines. Real-robot deployment further validates the feasibility of our method in real-world scenarios. More details are available on our project page: https://nav-ps-balance.github.io/.

Yifu Huo, Shunjie Xing, Chenglong Wang, ... , Tong Xiao, Jingbo Zhu
总结: 智能体RL用环境反馈实现多时间尺度信用分配。
方法: 长期结果信号辅以短期反馈与中期状态历史过程信号。
证据: 解决延迟稀疏奖励,提供细粒度中间决策监督。
为什么适合我: 信用分配改进长时域全身运动RL样本效率。
推荐理由: 面向长程智能体强化学习的信用分配,与人形全身控制、感知运动和动作重定向无关。
原摘要

Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. However, existing credit assignment approaches ignore the rich process information naturally generated during environment interaction, e.g., interaction history. We argue that such information provides valuable supervision for identifying the contribution of individual actions. To this end, we propose Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL. EFCA complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-term state-history signal that identifies ineffective patterns from recent interactions. Both signals are directly extracted from environment feedback and integrated through a return reweighting mechanism. Experiments on ALFWorld and WebShop demonstrate that EFCA consistently improves both task success and task quality over strong baselines, highlighting the effectiveness of environment-grounded multi-timescale credit assignment for long-horizon agentic RL.

Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
总结: 离线RL迭代贝尔曼残差选择可重用数据子集。
方法: 交替拟合批评家获取高残差转移后冻结静态子集。
证据: 10%预算保留96.6%性能,优于多数子集基线。
为什么适合我: 可重用离线数据选择支持高效全身运动策略训练。
推荐理由: 离线强化学习的数据子集选择,基准为通用D4RL,不涉及人形或腿足运动控制。
原摘要

Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment. We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing a reusable subset. Unlike prioritized replay, CODS produces a static artifact; unlike one-shot residual selection, it refreshes scores as the critic changes. At a 10\% budget, CODS retains 96.6\% of eligible-pool performance across 20 valid D4RL task--algorithm cells. It exceeds ReDOR and OPER on 19/20 cells and every other subset baseline on 20/20; all six subset advantages remain significant under predeclared hierarchical inference with Holm correction. Holding total selector updates fixed, five acquisition rounds improve four representative cells by 11.23 points over one round and saturate thereafter. Equal-pass and equal-hour evaluations clarify that reuse, rather than a single-run speedup, creates the compute advantage. Mechanism and corruption interventions expose both useful sparse-reward enrichment and sensitivity to outliers. Finally, a whole-trace extension retains 95.4\% of pooled ALFWorld success and 96.5\% of pooled GSM8K exact match. CODS is therefore a reusable selection procedure, not a formal coreset guarantee.

Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
总结: 平均奖励RL提供有限常数前沿与可审计遗憾证书。
方法: 常数感知比较协议推导通信MDP显式有限下界证书。
证据: 改进已发表系数,给出跨约束乐观学习可审计规则。
为什么适合我: 遗憾理论可指导人形机器人全身运动RL算法分析。
推荐理由: 平均奖励强化学习的遗憾下界理论,与机器人运动智能无关。
原摘要

Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ. We introduce a constant-aware comparison protocol and derive an explicit finite lower certificate for communicating MDPs. The construction is a binary tree of two-state blocks; its proof uses exact trajectory-level Bernoulli KL divergence and keeps action budget, diameter, occupancy, navigation cost, and terminal bias explicit. A common closed-form envelope improves the published coefficient $0.015$ across a finite frontier: $0.0200$ in a moderate regime and up to $0.0291$ under stronger action, diameter, and horizon conditions, a $94\%$ increase. The limiting coefficient is $\frac1{32}\sqrt{(A-3)/A}$. For upper bounds, we give an auditable composition rule for a span-constrained optimistic learner, but do not claim a coefficient while adaptive directional-variance and planning certificates remain open. We also formalize valid expectation conversion and constant comparability. Controlled diagnostics test diameter dependence, bonus-by-width interactions, span misspecification, and the finite lower certificate on its exact family.

Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han
总结: 流匹配模型用退化参考实现外推式在策略蒸馏。
方法: 隐式奖励外推转闭式速度回归,轻度退化强化对比。
证据: 桥接RL与OPD,实现稳定任务特定后训练。
为什么适合我: 流匹配蒸馏利于人体动作生成重定向到机器人控制。
推荐理由: 面向图像生成的flow-matching后训练,不是人体动作或机器人控制。
原摘要

Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.

The Sample Complexity of Policy Learning with Mu-Resets

1.0/5 偏低 裁判分 3.0 Candidate 当日相对
Gene Li
总结: μ重置协议下策略学习样本复杂度依赖覆盖假设。
方法: 分析策略可实现性,区分全策略与推前集中性覆盖。
证据: 全策略下指数下界,推前下紧致为exp(Θ(√H))。
为什么适合我: 样本复杂度理论优化人形机器人RL交互与探索协议。
推荐理由: μ-reset策略学习的样本复杂度理论,无全身运动或模仿控制内容。
原摘要

We study policy-based reinforcement learning under the $μ$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $μ$, in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon $H$ is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a $\exp(Ω(H))$ sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as $\exp(Θ(\sqrt H))$.

Distilling Physical Priors into Streaming World Models

1.0/5 偏低 裁判分 3.0 Candidate 当日相对
Liangliang Zhao, Junying Wang, Danni Yang, ... , Bowen Zhou, Yihao Liu
总结: 将物理先验蒸馏入流式世界模型保持长时物理一致。
方法: 三阶段框架建PhyS-120K物理交互数据集并监督微调。
证据: 克服双向教师先验有限及因果蒸馏损失问题。
为什么适合我: 物理世界模型增强复杂地形感知运动与跑酷预测。
推荐理由: 视频流式世界模型与物理视频先验,不落到可执行的人形/腿足控制;且包含软体等非兴趣现象。
原摘要

Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical constraints. A common approach distills pretrained bidirectional DiTs into few-step causal generators. However, this paradigm suffers from two fundamental limitations: generic bidirectional teachers acquire limited physical priors from visually oriented pretraining, and the limited priors suffer further loss during bidirectional-to-causal distillation. We present PhyS, a three-stage framework for distilling physical priors into streaming world models. To acquire physical priors from real-world interactions, we construct PhyS-120K, a dataset of 120K real-world physical-interaction videos spanning rigid-body dynamics, soft-body deformation, fluid phenomena, and phase transitions. Each video is annotated with structured descriptions of object properties and causal state transitions. Physics-aware supervised fine-tuning injects the physical priors into a bidirectional 14B DiT teacher, which we then distill into a lightweight 1.3B causal DiT for few-step autoregressive streaming generation. Finally, we use online reinforcement learning to incentivize the distilled model to generate physically plausible rollouts and further propose Temporal Credit Routing (TCR) to address temporal credit assignment. TCR evaluates physical consistency over overlapping temporal windows and routes the resulting group-relative advantages to temporally aligned denoising actions. On PhysicsIQ, PhyS improves the Wan2.1-14B teacher by 18.2\% and the Self Forcing, Rolling Forcing, and Causal Forcing by 23.7\%, 14.8\%, and 31.4\%, respectively. Results also improve the physics-aware video benchmarks VideoPhy, VideoPhy2, and PhyGenBench. The dataset, code, and more sample videos are available on our Project Page.