Papers for 2026-07-09

10 papers
Robert Jomar Malate, Erik Bauer, Norica Bacuieti, ... , Robert K. Katzschmann, Benedek Forrai
总结: 采样重定向实现低抖动实时手部运动映射。
方法: 提出梯度自由采样重定向器SBR用于运动学映射。
证据: 仿真与18人研究,成功率54.1%且疲劳最低。
为什么适合我: 手部重定向可迁移至人形全身动作真机控制。
推荐理由: 直接对应人体动作重定向与人形遥操作数据采集,用低抖动实时运动学重定向提升示教质量,虽限于手部仍值得看。
原摘要

Advances in learning-based robotic manipulation, such as Vision-Language-Action (VLA) models and Video Action Models (VAMs), heavily rely on high-quality teleoperation data. Their capabilities are strictly upper-bounded by the quality of the underlying human demonstrations. Current gradient-based retargeting algorithms often converge to different local minima, resulting in jitter that affects data quality and teleoperation experience. To address this, we introduce the Sampling-Based Retargeter (SBR), a novel gradient-free retargeting method drawn from the rich literature of sampling-based control and explicitly designed for low-jitter, real-time kinematic retargeting. We evaluate SBR both in simulation and through a rigorous real-world user study involving 18 participants performing 3 complex manipulation tasks. Compared to gradient-based baselines, SBR achieved the highest overall task success rate (54.1%) while significantly reducing operator cognitive fatigue, recording the lowest NASA-TLX workload score (36.4 out of 100). Ultimately, we establish SBR as a highly effective, intuitive retargeter for dexterous manipulation, providing the community with a rigorous benchmarking methodology to guide future retargeting research.

Sakuya Ota, Qing Yu, Kent Fujiwara, Satoshi Ikehata, Ikuro Sato
总结: 检索优化获胜噪声提升扩散运动语义一致性。
方法: 训练无关WINRO检索并KL优化初始噪声票。
证据: 改善文本运动对齐,支持组合长时序序列。
为什么适合我: 扩散人体运动生成利于动作重定向与跟踪。
推荐理由: 最接近扩散模型人体动作生成与运动先验,但只做文本-动作语义对齐,不涉及可跟踪、可重定向或上真机的全身控制。
原摘要

Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in compositional and long-duration sequences that require semantic consistency across multiple action segments and smooth kinematic transitions throughout the trajectory. We posit that the initial noise is central to this consistency: within the Gaussian noise space, certain instances, i.e. winning noise tickets, carry latent structure that biases denoising toward particular motion semantics, even under null prompts. We propose WInning Noise Retrieval and Optimization (WINRO), a training-free, model-agnostic framework that improves text-motion alignment by selecting and refining such tickets before diffusion sampling. WINRO maps random noises to motion features generated under null prompts, retrieves the best-aligned noise for a given text, and refines it via a KL-regularized objective that reduces the residual semantic gap while preserving the Gaussian prior. An optional LoRA-based adapter amortizes this refinement into a single forward pass. WINRO consistently improves text-motion fidelity across different base models, MDM and MotionLCM, on HumanML3D without retraining, improves temporal robustness on the MTT benchmark, and generalizes to applications such as motion stylization and spatial constraint satisfaction.

Ahmet Ercan Tekden, Yasemin Bekiroglu
总结: 物体中心神经场实现演示组合运动生成。
方法: 神经场表示加时序MoE组合运动原语。
证据: 仿真中实现跨多样场景配置系统泛化。
为什么适合我: 组合运动利于复杂接触场景操作学习。
推荐理由: 最接近模仿学习与场景交互的组合动作生成,但是通用示教运动基元,不是人形/腿足全身 loco-manipulation。
原摘要

Compositionality, by organizing complex behavior as combinations of simpler elements, enables robot learning that is scalable and data efficient. Leveraging this principle, we propose a generative learning-from-demonstration framework that enables compositional modeling of robotic behavior by connecting perception and motion through shared object-level representations. We render scenes from object-centric neural representations that integrate canonical neural fields with latent-conditioned deformations, capturing positional and geometric variations in a smooth, consistent, and interpretable way. For motion generation, a temporal mixture-of-experts (MoE) employs a gating mechanism to combine object-conditioned movement primitives over time, producing complete trajectories. This spatial-temporal compositionality maintains the data efficiency of movement primitives while grounding motion in visual structure, enabling systematic generalization across diverse scene configurations. In simulation, long-horizon manipulation tasks are successfully completed using the proposed model, which requires significantly less training data than other image-based baselines. Real-world experiments further demonstrate the method's robustness to noise, its ability to generalize at the category level through language-based segmentation models, and its capacity to operate directly on 3D scene representations.

Zezeng Li, Enda Xiang, Thuy Tran, ... , Momath Thiam, Liming Chen
总结: 测试时原语引导增强扩散策略自适应操作。
方法: PANet预测原语并微分引导精炼动作。
证据: 完全测试时运行,可无缝集成现有策略。
为什么适合我: 原语引导提升模仿学习全身控制泛化力。
推荐理由: 方法贴近扩散/流策略与模仿学习,但对象是通用操作策略自适应,而非全身运动跟踪或接触丰富运动。
原摘要

Imitation learning has enabled remarkable progress in robotic manipulation, especially with diffusion and flow-based policies that generate complex visuomotor behaviors directly from demonstrations. Yet, despite their strong performance, these policies often fail to generalize across tasks and environments. A key reason is that existing policies tend to imitate superficial action correlations rather than the underlying intent. Inspired by the compositional structure of human behaviors, we propose PriGo, a primitive-guided test-time adaptive framework for robust robotic manipulation. PriGo introduces PANet, a lightweight primitive prediction module that infers primitive distributions directly from observations. We further propose a differentiable primitive guidance mechanism that refines generated actions during inference, steering trajectories toward semantically consistent behaviors. Unlike prior primitive-conditioned approaches, PriGo operates entirely at test time and can be seamlessly integrated into pretrained diffusion and flow policies without retraining. Extensive experiments on LIBERO, CALVIN, SIMPLER, and real-world robotic tasks demonstrate that PriGo consistently improves robustness, long-horizon execution, and generalization ability across both diffusion and flow-based policies.

Residual-Conservative Model Predictive Path Integral Control

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Hyung-Jin Yoon, Hunmin Kim
总结: 残差保守MPPI在线调节安全应对模型失配。
方法: 残差约束收紧、安全成本与采样温度自适应。
证据: 推导约束违反概率界并验证联合效果。
为什么适合我: 采样MPC适用于腿足机器人全身运动控制。
推荐理由: 仅泛化采样式MPC与模型失配保守性,未落到全身模型预测控制或腿足接触控制。
原摘要

Sampling-based model predictive control methods handle nonlinear dynamics and complex cost landscapes through Monte Carlo rollouts, yet typically employ fixed constraint penalties that do not adapt to model-plant mismatch. This paper proposes Residual-Conservative Model Predictive Path Integral Control (RC-MPPI), a sampling-based MPC framework that modulates safety conservatism online using the prediction-execution residual. RC-MPPI combines three coupled mechanisms: residual-dependent constraint tightening, adaptive safety-cost shaping, and residual-adaptive sampling modulation through exploration contraction and temperature relaxation. The temperature adaptation reflects a key insight: when the model is inaccurate, rollout cost evaluations become unreliable, and increasing temperature reduces overcommitment to apparent cost rankings. Under Lipschitz dynamics and sub-Gaussian disturbances, we derive probabilistic bounds on constraint violation and show that the joint effect of the adaptive mechanisms reduces violation probability as the residual grows. A rollout-cost uncertainty analysis further shows that model-plant mismatch perturbs MPPI importance weights in proportion to residual magnitude and inversely with temperature, providing theoretical justification for residual-adaptive temperature relaxation. Simulations on an LTI point-mass system and a planar 2R manipulator show improved safety margin, success rate, and control efficiency compared with vanilla MPPI under significant model-plant mismatch.

EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Xinjie Wang, Liu Liu, Taojun Ding, ... , Wei Xu, Zhizhong Su
总结: 代理式3D世界引擎生成策略就绪仿真环境。
方法: 统一表示连接资产交互任务世界与编码。
证据: 资产接受96.5%、碰撞98.6%、83.3%直接可用。
为什么适合我: 生成环境支持人形复杂地形与操作训练。
推荐理由: 仿真世界资产生成可服务策略训练,但不是感知运动、跑酷或全身控制方法。
原摘要

We present EmbodiedGen V2, a generative 3D world engine for building executable policy-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapidly, yet assembling such assets into policy-ready task environments remains largely manual, limiting scalable closed-loop learning. EmbodiedGen V2 addresses this gap through a unified sim-ready representation that connects cross-simulator assets, interaction affordances, task-driven worlds, large-scale multi-room scenes, and stateful Vibe Coding into a generative, editable, and reusable simulation pipeline. The generated environments support manipulation, navigation, mobile manipulation, cross-simulator deployment, and embodied policy training. In evaluation, the asset pipeline achieves 96.5% human acceptance and 98.6% collision success, and 83.3% of task-driven worlds are directly usable for downstream simulation without manual modification. Online reinforcement learning with generated environments further improves simulation success from 9.7% to 79.8%, and transfers to real robots with task success increasing from 21.7% to 75.0%. These results establish EmbodiedGen V2 as scalable simulation infrastructure for training, evaluating, and deploying embodied policies.

Kaiwen Wang, Frank Bieder, Yinzhe Shen, ... , Jan-Hendrik Pauls, Omer Sahin Tas
总结: 小波相位扩散实现结构语义一致域翻译。
方法: 双树复小波包域操作避免全局频谱耦合。
证据: 消除振铃与边界泄漏等空间伪影问题。
为什么适合我: 域翻译利于仿真感知运动到真机迁移。
推荐理由: 只做视觉外观的sim-to-real翻译,不涉及动力学迁移或运动控制。
原摘要

Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism but rely on complex synthesis pipelines, often altering scene geometry and semantics. Training-free editing methods avoid both constraints but lack a learned appearance prior, limiting their perceptual quality. Recently proposed phase-preserving diffusion presents a promising alternative, but Fourier-domain formulations are constrained by global spectral coupling. This coupling induces spatial artifacts such as ringing and boundary leakage, thereby degrading structural and semantic consistency. We introduce Wavelet Phase Diffusion, which addresses this through two components. First, we operate in the Dual-Tree Complex Wavelet Packet Transform domain, whose localized wavelet packets enable spatially adaptive phase injection without global spectral interference. Second, Low-Frequency Randomization (LFR) replaces the low-frequency packet, decoupling the model from the synthetic illumination prior and enabling in-distribution real-world appearance. Both components train on unpaired open-domain data, and introduce negligible inference overhead. The spatial locality further enables instance-level translation, where individual objects or regions are translated to photorealistic appearance independently while the surrounding scene remains untranslated. On vKITTI $\to$ KITTI image translation, ours outperforms prior methods in realism and semantic consistency while maintaining competitive structural alignment. For CARLA video translation, ours approaches the realism of paired-data methods while reducing VLM planner ADE and FDE by $5.4\%$ and $5.1\%$, respectively.

Mathematical methods of reinforcement learning

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Denis Belomestny, Alexander Gasnikov, Egor Gladin, ... , Daniil Tiapkin, Nikita Yudin
总结: 综述强化学习概率优化与算子数学结构。
方法: 统一MDP贝尔曼、随机逼近与函数逼近。
证据: 覆盖收敛率、样本复杂度与约束RL等。
为什么适合我: 数学基础强化腿足机器人强化学习设计。
推荐理由: 通用强化学习数学综述,没有人形、腿足或运动跟踪内容。
原摘要

Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) and the Bellman operators, emphasizing contraction mappings, monotonicity, and fixed-point theory that yield convergence guarantees and rates for value and policy iteration, and temporal-difference schemes. We then develop the optimization perspective: stochastic approximation and martingale methods, convex duality and the role of regularization linking mirror/proximal methods. Function approximation is treated through linear and non-linear settings, covering stabilization, error decomposition, and sample-complexity via concentration inequalities for dependent data and mixing processes. We further cover off-policy evaluation/learning, constrained RL and constrained MDPs (CMDPs). Throughout we unify algorithmic templates under common operator and variational lenses, highlighting both finite-sample bounds and asymptotic results. Our presentation is intended to provide a unified mathematical entry point for researchers in probability, optimization, and statistics interested in reinforcement learning.

Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models

2.0/5 偏低 裁判分 2.0 Candidate 当日相对
Eli Laird, Corey Clark
总结: 解锁哈密顿视频动力学模型时间泛化能力。
方法: 修复动作力映射与积分器截断误差机制。
证据: 非保守环境中实现可变时间分辨率预测。
为什么适合我: 时间泛化世界模型利于复杂地形规划。
推荐理由: 视频动力学与时间尺度泛化,仅顺带提及sim-to-real,不是机器人运动控制。
原摘要

World models are typically trained to predict discrete-time physical dynamics with a fixed step size baked into the model weights, preventing prediction at variable temporal resolutions. This matters for hierarchical planning, sim-to-real transfer, and scientific or game-engine applications that must query the same dynamics at multiple timescales. Hamiltonian Generative Networks (HGN) offer a principled path forward, grounding predictions in a continuous-time energy function that is, in principle, independent of the observation frame rate. In practice, however, their temporal generalization breaks down in non-conservative settings. We show that in externally forced, dissipative environments, HGN rollouts at step sizes beyond the training regime fail due to distinct failure modes, including latent magnitude growth driven by an unconstrained action-force map, and global truncation error accumulation from an under-resolved integrator. We identify a targeted fix for each mechanism and demonstrate stable dynamics prediction at temporal resolutions well outside the training distribution. In a detailed analysis, we recommend several strategies for enabling temporal generalization in continuous-time video generation.

Initiation Safety: A Missing Dimension in Generalist-Robot Safety

2.0/5 偏低 裁判分 1.0 Candidate 当日相对
Zhijin Meng, Francisco Cruz
总结: 启动授权是通用机器人安全缺失的维度。
方法: 在人形上实现探测授权说话PAS框架。
证据: 对比直接启动日志并提出三条件研究。
为什么适合我: 启动安全完善人形接触丰富场景交互。
推荐理由: 虽出现门口人形,但讨论社交发起授权与安全,不是全身运动或loco-manipulation。
原摘要

Safety for generalist robots is usually discussed in terms of motion or dialogue. We argue a third question is missing: should the robot take its first hard-to-undo social action at all, such as a greeting, an uninvited grasp, or stepping into someone's space? We call this initiation authorization. Current frameworks rarely treat it as a separate safety layer. Today's stacks often skip this step: a high engagement score or a confident VLA rollout is treated as permission to act. But seeing a person is not the same as having their consent to be addressed. We frame initiation authorization within generalist-robot safety and contrast it with post-plan VLA guardrails, implementing PAS (probe-authorize-speak) on a doorway humanoid, comparing it with direct-init on logged traces, and proposing a three-condition user study, with open questions on metrics, governance, and where initiation ends and foundation-model generation begins.