Papers for 2026-09-07

10 papers

STyMo: Fast and Controllable Few-Shot Motion Style Transfer

3.0/5 一般 裁判分 3.0 Candidate 当日相对
Jose Luis Ponton, Alexander Winkler, Ladislav Kavan, Yuting Ye, Petr Kadlecek
总结: 少样本运动风格迁移,秒级数据分钟级训练可控。
方法: 分解风格为静态姿态与时序动态并加门控。
证据: 支持运行时调节与分布外输入鲁棒演示。
为什么适合我: 风格迁移可助人体动作重定向全身控制。
推荐理由: 最接近物理角色控制与运动先验:少样本运动风格分解与体段可控,但是虚拟角色运动风格编辑,不是可跟踪、可重定向的全身机器人控制。
原摘要

Supporting a wide variety of motion styles is critical for creating diverse virtual characters, but current methods either require large stylized datasets or pre-trained models that cannot generalize beyond their training distribution. We present STyMo, a few-shot approach that learns motion style from only seconds of paired data and trains in one to two minutes. Our key insight is to decompose style into two components: a static component capturing time-invariant posture, and a temporal component capturing frame-wise dynamics. This decomposition yields an interpretable system where posture intensity, temporal exaggeration, and per-body-region style can be adjusted at runtime. Furthermore, the reduction in required training data and computation time structurally permits an iterative authoring workflow. To ensure robustness on arbitrary inputs, we further introduce a stylizability gate that automatically prevents artifacts on out-of-distribution motions. We demonstrate results across diverse motion styles, from subtle emotional variations to exaggerated character archetypes, and release our processed paired dataset to facilitate future research.

Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Amarjot Singh, Tanmay R. Pancholi, Jainam Kothari, ... , Jeff Schneider, Vince Nakayama
总结: 部署后持续自适应物理AI,支持现场无梯度更新。
方法: 慢学习皮层加快速胶囊场互补架构。
证据: 实验室少样本与现场持续学习技能安装。
为什么适合我: 利于腿足机器人复杂地形部署后适应。
推荐理由: 最接近部署后物理智能与技能执行,但是通用持续学习架构,未涉及人形/腿足全身控制、感知运动或真机运动跟踪。
原摘要

Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations. These domains offer scarce training data and only onboard compute, yet deployed systems must face novelty without erasing prior competence. We introduce Continual Field-Adaptive Models (CFAMs), which learn efficiently in the lab and continue learning after deployment through autonomous, gradient-free, on-device updates. CFAM uses a complementary learning architecture with a frozen slow-learning component and a fast-learning Capsule Field. The slow component contains three cortices: Sensor, which maps multimodal input into 3D-grounded geometry; Reasoning, which decomposes tasks into skills and evaluates outcomes; and Action, which executes geometric skills. The Capsule Field stores field learning one-shot and gradient-free as Competence Capsules. Skill installation is few-shot in the lab and continual in the field; open-world novelty is outside scope. We evaluate CFAM across five embodiments: manipulator, quadruped, humanoid, quadrotor, and off-road vehicle. Baselines (pi0, CogACT, SpatialVLA) use the same in-house multi-embodiment dataset for physical-platform comparisons. CFAM reaches the operating point of a standard policy trained on the full prior-training dataset using 40% of the data, or 2.5x fewer trajectories. At test time, autonomous capture of verified near-OOD cases improves action success by 13.9 percentage points. In sequential simulation, backward transfer is -0.5 percentage points versus -11.4 for LoRA. CFAM therefore provides a bounded form of post-deployment physical intelligence: few-shot skill learning, autonomous field growth from verified near-OOD experience, and retention of prior competence.

Yirong Zeng, Shen You, Jinhang Feng, ... , Wang Xu, Bibo Cai
总结: 自动合成可执行环境供爪型智能体强化学习。
方法: 环境合成引擎与拓扑感知轨迹生成。
证据: 合成139环境约两万复杂任务用于训练。
为什么适合我: 偏LLM智能体,与全身运动关联较弱。
推荐理由: 面向LLM智能体的可执行环境合成与Agentic RL,与人形腿足运动无关。
原摘要

The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its scaling is heavily bottlenecked by the severe scarcity of interactive training environments. Existing synthetic environments are strictly limited to tool-calling endpoints, rendering them insufficient for accommodating the end-to-end real-world demands of claw-like agents. To bridge this gap, we introduce EnvCraft, an automated framework for synthesizing executable environments and scalable training data. Specifically, EnvCraft employs an environment synthesis engine to build sandbox-isolated workspaces, alongside a topology-aware data generation engine to produce coherent task trajectories. Overall, we synthesize 139 interactive environments comprising approximately 20K complex tasks for Agentic RL training. Experiments on Qwen3/3.5 models (8B-32B) show that our method yields gains of up to +11.9% on Claw-style benchmarks and +8.0% on general tool-use benchmarks, with concurrent reductions in inference token cost. The results confirm that synthesized executable environments provide robust and generalizable learning signals for training.

Marcel Moll, Timo Oksanen
总结: 拖拉机多参考路径跟踪的非线性模型预测控制。
方法: 多段分段线性路径纳入目标的NMPC。
证据: 田间测试求解3.45毫秒横向误差6.1厘米。
为什么适合我: NMPC可借鉴腿足机器人全身MPC控制。
推荐理由: 农用拖拉机路径跟踪NMPC,属于车辆轨迹控制,不是腿足或人形全身MPC。
原摘要

Guiding a tractor along a predefined reference path is a key component of precision agriculture. This study develops a path tracking controller based on Nonlinear Model Predictive Control, which incorporates multiple segments of a piecewise-linear reference path directly into the objective function. In addition, methods for selecting viable reference segments from the full path are presented. The control system is evaluated during a field test with a tractor controlled via the Tractor Implement Management steering interface. The NMPC solver converged on average after 3.45 ms and tracked the curved reference path with a mean absolute cross-track error of 6.1 cm.

Jinyuan Feng, Dongmin Li, Yiqun Chen, ... , Huimu Wang, Zhiqiang Pu
总结: 联合强化学习推理与元技能实现层次技能进化。
方法: 元技能建模为可学习智能体并联合训练。
证据: 端到端共适应促进技能灵活进化。
为什么适合我: 层次技能RL启发loco-manipulation技能库。
推荐理由: LLM推理与元技能智能体的联合强化学习,不涉及机器人运动控制。
原摘要

Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large language model (LLM) agents to reuse procedural knowledge. Yet existing paradigms exhibit structural shortcomings: they either decouple skill evolution from policy optimization or instantiate meta-skills as fixed workflows. Both treat skills as passive objects to be managed, limiting the flexible evolution of skills and their co-adaptation with the reasoning agent. To address the limitations, we propose CoSkill, a unified multi-agent RL framework that recasts the static meta-skill workflow as a learnable Meta-Skill Agent and jointly trains it with a Reasoning Agent over a hierarchical skill library. By modeling the Reasoning and Meta-Skill Agents as a cooperative team sharing a single backbone, CoSkill enables end-to-end co-adaptation: the Reasoning Agent conditions its actions on a retrieved task skill and step skills selected from its child set, while its task performance guides the Meta-Skill Agent in refining those step skills. Experiments on ALFWorld and WebShop show that CoSkill substantially outperforms prior skill-based and RL baselines, achieving success rates of 98.4% and 90.6%, respectively (+3.5 and +6.2 pp). As shown in Figure 1, CoSkill achieves superior early-stage sample efficiency, asymptotic performance, and wall-clock efficiency. Our code is available at https://github.com/jinyuan-cookie/CoSkill.

Heeirthan Shanthan, Winston Hurst, Yasamin Mostofi
总结: 动态信道下机器人运动通信能量联合优化。
方法: NMPC框架加控制屏障函数协同决策。
证据: 动态障碍下安全导航与及时数据传输。
为什么适合我: 动态环境NMPC利于复杂地形接触控制。
推荐理由: 自动驾驶车辆在动态障碍与毫米波链路下的运动—通信联合NMPC,落在不感兴趣的自动驾驶方向。
原摘要

This paper studies energy-efficient operation of autonomous vehicles (AVs) in dynamic environments with moving obstacles and while communicating over mmWave channels. The obstacles induce severe attenuation of the mmWave channel resulting in a highly dynamic communication environment. In this setting, we consider the problem of jointly optimizing motion and communication energy for an AV that safely navigates among dynamic obstacles toward a designated destination while ensuring timely transmission of onboard sensing or telemetry data over mmWave channels. We then seek a real-time methodology to compute energy-efficient trajectories in a setting where dynamic obstacles induce both safety constraints and time-varying mmWave blockage, leading to tightly coupled motion-communication trade-offs. We propose a nonlinear model predictive control (NMPC) framework that enables anticipative communication and motion decision-making and energy co-optimization, augmented with a control barrier function (CBF) to ensure safety. Extensive simulation results demonstrate the effectiveness of our approach, reducing total energy consumption by up to 37.3% compared to baseline strategies. Overall, our results demonstrate that the proposed NMPC-based framework significantly enhances energy efficiency and performance of AVs under dynamic, blockage-sensitive mmWave communication constraints.

Compositional Reward Models for Conditional Medical Image Generation

1.0/5 偏低 裁判分 1.0 Candidate 当日相对
Aayush Kumar Tyagi, Prathosh A. P., Mausam
总结: 组合奖励模型提升条件医学图像生成质量。
方法: 分解质量为多阶段验证器接地奖励。
证据: 避免单一标量奖励混淆失败模式。
为什么适合我: 医学图像生成与机器人运动无关。
推荐理由: 条件医学图像生成的组合奖励模型,与运动智能无关。
原摘要

Acquiring high quality annotated medical image data is critical for training deep learning models; however, annotation is expensive, time consuming, and requires domain expertise. Conditional diffusion models, such as ControlNet, offer an alternative by generating images conditioned on semantic masks and text. However, existing approaches fail to capture fine grained properties (e.g., intensity and texture), as well as semantic consistency expected by domain experts, limiting their effectiveness for downstream tasks. Recent attempts to address these issues using reinforcement learning fine-tuning remain limited due to the reliance on a single scalar reward, which conflates diverse failure modes and provides weak corrective signals. We propose PRISM, a Compositional Reward Model (CRM) framework for conditional medical image generation. Instead of assigning a single reward, we decompose image quality into verifier grounded stages, each evaluating a distinct aspect of correctness from fine to coarse properties, including low level attributes (intensity and texture), structural alignment with conditioning inputs, and high level semantic fidelity. These stage wise rewards are composed through a Hierarchical Constrained Propagation (HCP) mechanism that enforces a fine to coarse notion of correctness, ensuring that lower level deficiencies are resolved before higher level rewards are accrued, preventing easier objectives from masking critical failures. We evaluate PRISM across three datasets spanning diverse medical imaging tasks: PanNuke (multi-class cell segmentation), CeDeM (villi/crypt detection and measurement), and ISIC (skin lesion classification). Training downstream models with data generated by PRISM yields improvements over closest baselines, including a 2.3% increase in mDice on PanNuke, a 8.5% reduction in Mean Relative Error (MRE) on CeDeM, and increases ISIC F1 by 5.9%.

Maria Mahbub, Ashley Rice, Michael R. Munroe, Amidu Kamara, Amir Sadovnik
总结: 行为正确性假设评估参考式自动评估方法。
方法: 定义保持与改变假设经受控变换验证。
证据: 揭示评估器权衡无一满足全部假设。
为什么适合我: NLP评估与人形腿足运动控制无关。
推荐理由: 自然语言生成的自动评测行为假设,与机器人控制无关。
原摘要

Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Our experiments reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can exhibit substantially different behavioral profiles. These findings demonstrate that behavioral correctness assumptions provide diagnostic information obscured by conventional aggregate meta-evaluation.

Chen Cheng, Michael Ferraro, James Grover, David E J Waddington, Emily Hewson
总结: 束眼视图Mamba-3架构实现隐式剂量重建。
方法: 状态空间深度核心加物理传输条件。
证据: 支持光子质子剂量快速准确计算。
为什么适合我: 医学剂量计算与全身运动智能无关。
推荐理由: 放疗剂量重建的束眼视图模型,不属于机器人感知或运动。
原摘要

To enable accurate and rapid photon control point and proton beamlet dose calculation in the DoseRAD2026 challenge, we present BEAM3R, a dose estimation framework operating in beam's-eye-view (BEV). Our core innovation combines a Mamba-3 state-space depth-sequence core with physics-based transport conditioning to model long-range depth transport without expensive 3D convolutions. BEAM3R shares a 2D CNN encoder-decoder architecture for photon and proton dose tasks, processing per-plane BEV slices. Proton beamlets are conditioned on water equivalent thickness and remaining range, encoding the parameters determining Bragg peak position. Photon models use a bidirectional Mamba-3 core to capture dose contributions from materials downstream of the calculation point, while the proton model uses a forward core with learned energy-prefix tokens and a Bragg-peak refinement module. To reduce interpolation artifacts and support high spatial resolution, we introduce axial grid alignment of BEV lattices with CT slices and an implicit super-resolution representation via sub-pixel phase packing, evaluated by a differentiable Triton-accelerated resampler that reconstructs packed cubic B-spline coefficients directly in CT space. For MRI-based tasks, synthetic CTs (sCT) are generated by a patch-based conditional GAN with a SwinUNETR backbone. On the preliminary DoseRAD2026 test set, CT-to-photon and CT-to-proton models achieved 1%/1 mm local gamma pass rates of 96.8% and 96.0%, with stratified plan-level MAEs of 0.0041 and 0.0079. Substituting sCT reduced gamma pass rates to 89.7% for photon and 75.4% proton plan level doses, with stratified plan-level MAEs of 0.0093 and 0.0336. Standardised runtimes were 23.4 s and 18.4 s for CT-to-photon and CT-to-proton prediction, increasing to 39.7 s and 42.8 s for the corresponding MRI-based pipelines.

A Sim-to-Real Study of Surface-Code Decoder Benchmarking

1.0/5 偏低 裁判分 0.0 Candidate 当日相对
Shay J. Manor, Leila S. Erhili, Yassine Jebbouri
总结: 表面码解码器仿真到真实基准排名研究。
方法: 递增保真噪声阶梯评估解码器一致性。
证据: 特定噪声后排名一致并评估预解码器。
为什么适合我: 量子纠错与腿足机器人运动无关。
推荐理由: 量子表面码译码器的sim-to-real基准,仅名词相近,不是机器人仿真到真机。
原摘要

Quantum error-correction decoders are typically benchmarked against synthetic circuit-level noise, under the assumption that a decoder's ranking under such noise transfers to hardware and improves as the noise model becomes more realistic. The Willow processor, the first to operate below the surface-code threshold, allows us to test this assumption. We rank a panel of six decoders using a four-rung ladder of noise models with increasing fidelity, evaluated against real data across three code distances, two bases, and fifteen round counts. Rank agreement with hardware appears once the noise model gives each operation type its own error rate. Calibrating the model to the device improves absolute error rates but not rank agreement. We additionally provide the first independent evaluation of NVIDIA's Ising pre-decoder on hardware, at code distances below its training receptive field and via a mapping onto the lattice on which it was trained. Under these conditions, it holds no accuracy-latency advantage: another panel decoder matches or improves on it in both per-cycle error rate and decode latency in 278 of the 280 evaluations. We release the full pipeline and the per-shot outcome of every evaluation, so future decoders and devices can be compared.