Papers for 2026-09-24

10 papers

Banana Kick: Response-Informed Skill Evolution for Humanoid Soccer

5.0/5 很相关 裁判分 9.0 Top pick 当日相对
Hao E. Zhang, Ruize Geng, Raihan Haque, ... , H. Eric Tseng, Ding Zhao
总结: 提出RISE解决人形香蕉球踢球先验适应中的一阶学习饥饿。
方法: 闭环目标延续,用缓存rollout响应敏感度排序并验证更新。
证据: 解决局部平坦目标导致的学习饥饿,保持踢球可靠性。
为什么适合我: 适用于接触丰富人形全身踢球技能的强化学习进化。
推荐理由: 直接对应人形全身控制与接触丰富技能:以模仿得到的踢球先验为起点,用强化学习把技能演化到香蕉球这类不同接触力学,属于运动先验上的策略适应。
原摘要

Humanoid kicking requires coordinated whole-body motion and precise contact, while a banana kick demands contact mechanics that generate ball spin and aerodynamic curvature. Motion imitation provides a reliable ordinary-kick prior, but reinforcement learning may improve shot speed and placement accuracy without changing the underlying kicking technique. Adapting this prior to a qualitatively different contact-rich skill can fail even when the reward is dense and optimization remains stable. The failure occurs when the task objective is locally flat over the current policy's responses. We term this condition first-order learning starvation. To address it, we propose response-informed skill evolution (RISE), a closed-loop objective-continuation method for policy adaptation. RISE ranks bounded objective changes using response sensitivity estimated from cached rollouts and accepts updates only when they produce verified response progress while preserving kicking reliability. Our analysis shows that rescaling a saturated spin reward cannot recover first-order sensitivity at zero spin, whereas adapting coupled contact responses can provide a learnable path to spin generation. We integrate RISE into a humanoid kicking pipeline under calibrated contact and Magnus-force aerodynamics, and sim-to-real transfer. Experiments show that RISE evolves the ordinary kick into a high-spin curved kick with 11.55 rad/s mean ball spin, improves the mean evaluation score by 19.8% over a learning-progress curriculum, and raises joint target attainment from 15.2% to 50.9%. Ablations and response diagnostics support the mechanism, while 30 motion-capture-recorded physical trials demonstrate consistent hardware transfer of the learned curved kick. Project website: https://haozhang-thu.github.io/bananakick/

Jiakang Jin, Yixiao Huo, Pengyuan Wang, ... , Xiaoyu Tian, Yiming Li
总结: 仅用头部深度图端到端学习人形足球接触技能的框架。
方法: 学可见性辅助几何,结合GT退火、任务课程和AMP先验。
证据: 直接输出25自由度关节PD目标,无需额外模块。
为什么适合我: 深度感知与AMP契合人形接触丰富场景全身控制。
推荐理由: 最贴近腿足感知运动与接触技能:仅用头部深度和本体感觉端到端输出人形关节目标,并结合AMP式运动先验完成足球接触、对准与恢复。
原摘要

Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close the loop over perception, approach, alignment, impact, and recovery while its own motion induces substantial viewpoint changes, frequent loss of the ball from view, and uncertain contact outcomes. In this work, we ask a compact yet stricter question: can a humanoid learn soccer contact skills using only a head-mounted depth image, proprioceptive history, and an optional low-dimensional task command, and directly output 25-DoF joint PD targets without extra runtime perception or planning modules? To this end, we propose DAVIS, a depth-only end-to-end framework for humanoid soccer skills that learns visibility-aware auxiliary geometry during training, and combines GT-to-prediction annealing, task curricula, and AMP-style motion priors to smoothly bridge privileged supervision and real deployment. Built on this framework, we instantiate representative soccer contact skills, including goal-directed shooting and directional dribbling, through task-specific definitions of objects, commands, rewards, and curricula, and validate them through simulation, Noetix E1 real-robot experiments, and ablations.

Humanoid Locomotion with a Fly-Inspired Recurrent Controller

3.0/5 一般 裁判分 9.0 Top pick 当日相对
Isabel Guan, Yuntian Zhao, Dingyuan Zhang, Shipeng Lyu
总结: 苍蝇启发循环控制器支持人形机器人运动部署行为。
方法: 耦合神经状态与身体观测投影、运动神经元读出和伺服。
证据: 多地形速度偏航下多数成功,重置状态导致失败。
为什么适合我: 启发人形腿足全身运动的循环神经控制方法。
推荐理由: 对象是Unitree G1人形在多种地形上的运动,贴近腿足运动,但核心是果蝇启发循环控制器的通路分析,而非强化学习、模仿、对抗运动先验或全身MPC。
原摘要

We investigate humanoid locomotion with a fly-inspired recurrent controller and identify the pathways supporting its deployed behavior. The controller couples 3,609 continuous neural states to a simulated Unitree G1 through body-observation projections, a motor-neuron-labelled readout, and joint servos. We formulate this neural-body feedback system and evaluate a fixed checkpoint across seven terrain instances, three speeds, and three initial yaw offsets. It completes 61/63 conditions under a survival-and-forward-progress criterion; a privileged reference completes 62/63. At nominal yaw, resetting the recurrent motor state before every policy call changes success from 19/21 to 0/21. Conversely, depth and upstream-state substitutions at 252 recorded states leave actions unchanged, with zero measured descending output throughout the intact rollouts. Recorded trajectories and state-matched images connect these findings to sustained movement, lateral drift, and termination events. The study characterizes an embodied recurrent control system whose tested locomotion is supported by direct body-and-command input and carried motor state, providing a concrete basis for subsequent comparisons of circuit structure and control resources.

Chen Xu, Rishi Shah, Hadas Kress-Gazit, Haruki Nishimura, Masha Itkina
总结: 微调大型行为模型时非高斯流匹配先验并无帮助。
方法: 比较高斯与接近目标非高斯先验在LBM微调中的效果。
证据: 大量仿真和硬件rollout显示无差异或更差。
为什么适合我: 指导扩散模型在模仿学习全身控制微调的先验选择。
推荐理由: 方法触及流匹配与扩散式生成策略,但问题是通用大行为模型的操作微调,不是人形全身跟踪、跑酷或loco-manipulation。
原摘要

Modern robot imitation learning increasingly relies on generative policies based on diffusion or flow-matching models, which generate actions by transforming samples from a prior distribution. A key question is whether the choice of prior matters. Replacing the standard Gaussian with a closer-to-target, non-Gaussian prior has been shown to substantially improve performance when training from scratch. A natural next step is to ask whether these gains transfer to fine-tuning pretrained Large Behavior Models (LBMs) such as LBM 1.0, $π_{0.5}$, and GR00T~N1.5, where one might expect even larger gains. Surprisingly, we find that this is not the case, except possibly at very low fine-tuning data fractions. Across over 100K simulation rollouts spanning all three aforementioned LBMs on 40+ tasks in two simulation platforms, and 1250 hardware rollouts on five bimanual manipulation tasks, non-Gaussian priors that are demonstrably closer to the target yield statistically indistinguishable or worse fine-tuning performance than a standard Gaussian prior. Diagnostic analyses suggest why: fine-tuned imitation learning policies converge to similar action predictions across priors, despite their fine-tuned encoder embeddings diverging substantially from the pretrained embeddings and each other. A learning-rate ablation further confirms that encoder training is the dominant factor in fine-tuning performance, substantially outweighing the effect of prior choice. We conclude with concrete directions for future research on when and why learned priors might still matter in fine-tuning. Project page: https://cxu-tri.github.io/non_gaussian_FT/

Weihui Zhao, Xiaohan Yan, Zunian Wan, ... , Wei Shan, Maoqing Yao
总结: 干预自适应真实世界RL框架让VLA超越专家模仿。
方法: 纠正模型预测人类纠正方式及各动作维度一致性。
证据: 将纠正视为约束证据,优化精度关键相位动作。
为什么适合我: 支持接触丰富场景VLA策略真实机强化学习。
推荐理由: 属于真实机器人上的VLA在线强化学习与人工干预,聚焦通用精细操作,不涉及人形全身运动或腿足接触控制。
原摘要

Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. However, existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Yet human corrections are not uniformly noisy but reliable along some action dimensions and variable along others. Building on this, we introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. We formulate human corrections not as actions to reproduce but as evidence about a constraint: a Correction Model predicts how a human would correct a given VLA proposal and how consistent the correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human, and where they vary, the constraint relaxes. We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.

HiRE: Hindsight Reward Editing for Policy Finetuning

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Haoyi Niu, Zhengtao Han, Yufeng Ji, Zhongyu Li, Koushil Sreenath
总结: 事后奖励编辑框架突破预训练策略微调奖励瓶颈。
方法: 对比成功失败轨迹识别陷阱状态并编辑奖励。
证据: 无需训练校准基础模型与物理控制意识。
为什么适合我: 提升人形全身策略强化学习微调的奖励质量。
推荐理由: 讨论预训练机器人策略的奖励微调,属通用强化学习工具,未落到人形全身控制、动作重定向或感知运动。
原摘要

Pre-trained robot policies always require finetuning to adapt to specific environments. Reinforcement Learning (RL) offers high performance potential because it improves action optimality rather than simply mimicking data. However, such potential depends heavily on reward quality. Sparse rewards lack process feedback, human-designed rewards are costly and biased, and semantic rewards from foundation representations are often not control-centric. We propose Hindsight Reward Editing (HiRE), a training-free framework to break this reward bottleneck. HiRE bridges the broad knowledge of foundation representation models with physical control awareness, by contrasting successful and failed trajectories in hindsight. It calibrates foundation representation models by identifying "trap states" that are predicted as high-rewarding states yet eventually result in failure, and vice versa. HiRE explicitly penalizes these traps while boosting rewards for critical successful states. This approach can be flexibly compatible with any foundation representations and RL algorithms. Experiments show that HiRE consistently outperforms other reward recipes by delivering dense, control-aware feedback that prevents value function collapse and reward hacking, thereby achieving superior sample efficiency, stable policy updates, and higher performance ceilings, e.g., at least 3x performance of the base policies. Qualitative results are at https://hire-project.github.io .

X2Real: an eXtensive simulation benchmark for real-world generalist policies

2.0/5 偏低 裁判分 5.0 Candidate 当日相对
Lian Ruan, Jade Yang, Sherphylan Gao, ... , Hao Wang, Qian Wang
总结: 可进化仿真基准忠实评估真实世界通用操作策略。
方法: 校准仿真视觉物理对齐硬件,含十能力维度任务。
证据: 仿真真实评估线性相关0.84,覆盖四十四层次任务。
为什么适合我: 为loco-manipulation策略提供高保真公平评估基准。
推荐理由: 是面向通用操作策略的仿真评测基准,虽涉及sim-to-real,但不是人形或腿足全身运动智能。
原摘要

Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.

Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu
总结: 两阶段安全滤波RL学习最小可利用竞争机器人策略。
方法: 对抗RL学安全滤波器,嵌入训练并部署保留。
证据: 触地游戏胜率Elo最高可利用性最低,硬件测试有效。
为什么适合我: 增强人形竞争任务中安全全身运动控制。
推荐理由: 安全过滤强化学习用于对抗性竞争任务,场景是触地类博弈而非人形全身运动、跑酷或场景交互。
原摘要

Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.

Latent evolving World Action Model

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Xueji Fang, Boqiang Duan, Hua Wu, Jingdong Wang, Guo-Jun Qi
总结: 潜在演化世界动作模型用JEPA嵌入改进动作生成。
方法: 条件动作于JEPA嵌入,预测未来嵌入建模演化。
证据: JEPA预测嵌入优于VAE压缩潜在支持动作生成。
为什么适合我: 改进世界模型助力人形感知运动与loco-manipulation。
推荐理由: 世界动作模型与视频表征用于动作生成,属于通用视觉-动作建模,而非人形运动跟踪或loco-manipulation。
原摘要

World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations. With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28\% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.

Dissecting Advantage-Guided Post-Training for Vision-Language-Action Policies

2.0/5 偏低 裁判分 3.0 Candidate 当日相对
Jiahang Cao, Hanye Zhao, Hang Lai, ... , Yong Yu, Weinan Zhang
总结: 剖析优势引导VLA后训练中设计选择的独立效应。
方法: 分离优势构建校准加权,用阶段离线评估筛选配方。
证据: 模块化配方在四个真实双臂操作任务中有效。
为什么适合我: 优化优势引导RL用于人形VLA全身控制后训练。
推荐理由: 拆解VLA策略的优势引导后训练,对象是视觉语言动作操作策略,不是人形全身控制或人体动作重定向。
原摘要

Advantage-guided reinforcement learning provides a practical way to post-train vision-language-action (VLA) policies using limited robot data. However, its performance depends on several coupled choices, including how critic-derived advantages are constructed, calibrated, and used for policy training. Existing recipes often combine these choices into a single end-to-end procedure, making their individual effects difficult to identify. In this work, we dissect advantage-guided VLA post-training through a controlled empirical study that separates these design choices while accounting for their distinct estimands. We develop stage-specific offline evaluation methods to screen alternative choices efficiently, without requiring extensive real-robot policy evaluations for every possible combination. The staged evaluation identifies a modular recipe that combines temporal-difference advantage construction, group-wise calibration, and continuous advantage weighting. Across four real-world bimanual tasks, the resulting recipe improves mean task progress and success over the SFT initialization by 0.42 and 0.63, respectively. Moreover, the proposed evaluation diagnostics show an overall alignment with downstream real-world performance, supporting their use for interpreting empirical outcomes and selecting advantage-guided post-training designs in practice.