Papers for 2026-08-25

10 papers
Aniruddh Kushwah, Vyankatesh Ashtekar, Ashish Dutta
总结: 参考引导强化学习为人形机器人生成可变形地面站立动作。
方法: 用MuJoCo软接触模拟地形,残差控制跟踪演示并加恢复奖励,先硬后软训练。
证据: 策略成功完成跌倒站立,适应穿透延迟力并保留演示模式。
为什么适合我: 高度契合人形接触丰富环境全身控制与RL参考跟踪迁移。
原摘要

This paper presents a reference-guided reinforcement learning framework to generate stand-up motion for a 29-DOF Unitree G1 humanoid on deformable soft ground, using a human demonstration recorded on hard ground. The terrain compliance is modelled using solref and solimp parameters from MuJoCo's rigid body soft-contact model. The rewards consists of (i) reference motion tracking through residual joint-position control and (ii) explicit recovery objectives such as pelvis height, torso uprightness, and the final posture. First, the policy is trained with the specified rewards considering hard ground. Next, the terrain stiffness is lowered by updating solref and the nominal surface penetration zone is expanded using solimp. Subsequent training enables the policy to adapt to the delayed support force generation due to significant surface penetration during contact-intensive phases while preserving the original demonstration pattern. The learned policy successfully completes the fallen-to-standing task in simulation, reaching the targeted pelvis height and uprightness, with a maximum contact penetration of approximately 40 mm during the process. The proposed method is demonstrated on two stand-up sequences and successfully achieves the final recovery objective on both hard and soft ground. Ablation studies show that reference tracking alone is insufficient for successful stand-up, and that explicit recovery rewards are essential.

Giray Onur, Azita Dabiri, Bart De Schutter
总结: 提出DRL与MPC共享控制权威框架用于多类交通网络。
方法: 划分控制权限结合DRL学习与MPC优化应对高维非线性系统。
证据: 缓解DRL容量限制与MPC计算负担,利于实时大规模控制。
为什么适合我: 交通网络控制与腿式人形机器人全身运动无关。
原摘要

Transportation networks, in particular multi-class transportation networks (i.e., networks with mixed vehicle types), are complex systems that are challenging to control. Recently, Deep Reinforcement Learning (DRL), which learns control policies from interactions with the environment, and Model Predictive Control (MPC), which uses a system model to optimize control inputs, have been increasingly utilized for transportation network control. However, nonlinear system dynamics and high-dimensional state spaces in large-scale networks limit DRL's learning capacity under time-constrained training and increase MPC's computation time, hindering real-time implementation with limited computational resources. Moreover, MPC depends on an accurate network model, which is often unavailable for complex systems such as multi-class transportation networks. This paper proposes a novel DRL-MPC framework for multi-class transportation networks that divides control authority between DRL and MPC, combining DRL's fast online computation and model independence with MPC's built-in optimization and constraint-handling capabilities. In the hierarchical framework, MPC operates at the higher level and determines low-frequency control inputs whose slower update rate accommodates its high computation time, while DRL operates at the lower level and determines high-frequency control inputs using its fast online deployment. The framework is evaluated on a multi-class freeway network against a hierarchical MPC controller and a hybrid state-feedback-MPC controller, including scenarios with model mismatch and noisy traffic demands. Results show that the proposed framework outperforms the hybrid state-feedback-MPC controller, substantially reduces online computation time compared with the hierarchical MPC controller, and provides more effective constraint enforcement under model mismatch.

Vincent K. M. Cheung, Jia-Yeu Lin
总结: 提议人形音乐机器人作为音乐诱发情感实验接口。
方法: 参数化控制表演变量,实现听觉视觉交互成分解耦重组。
证据: 展示技术可行性,超越预录刺激捕捉多模态体验。
为什么适合我: 人形机器人用于情感研究,非接触运动控制核心。
原摘要

Advances in technology have led to increasingly sophisticated musical humanoid robots. However, their use has largely been limited to performance and related research in human-robot interaction. In this position paper, we propose a novel perspective: musical humanoid robots as experimental interfaces for investigating music-evoked emotions. We argue that current research is constrained by paradigms relying on pre-recorded auditory stimuli, which fail to capture the multimodal, embodied, and interactive nature of real-world musical experience. Building on existing theories of music cognition and emotion, we identify mechanisms that require controlled manipulation of both acoustic and non-acoustic variables. We show that humanoid robots are well-suited as they enable parametric control of performance variables, reproducibility across trials, and the decoupling and recombination of auditory, visual, and interactive components. We illustrate the technical feasibility of this perspective through a case study of the WAseda Saxophonist Robot 5 (WAS-5), demonstrating reproducible control of acoustic and interaction variables that are prerequisites for future music-emotion experiments. Our work positions musical humanoid robots as a methodological platform that enables future controlled investigations of music-evoked emotions.

Zhitao Liu, Guangtong Xu, Zihan Wang, ... , Chao Xu, Fei Gao
总结: 提出原语模仿学习高效端到端局部规划器用于自主飞行。
方法: 离线原语数据集训练紧凑网络,直接映射感官到多项式系数。
证据: 实时生成平滑可行轨迹,计算超快内存低,零样本真实部署。
为什么适合我: 飞行器规划方法,与腿式人形全身控制无关。
原摘要

Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this paper, we propose an efficient end-to-end local planner via imitation learning. A lightweight offline-primitive-based dataset collection framework is designed to produce safe and high-quality trajectory primitives in non-convex environments. A compact neural network directly maps sensory inputs to polynomial coefficients that inherently encode higher-order dynamical information. The learned policy generates smooth, empirically collision-free and dynamically feasible trajectories in real time without back-end solving. It achieves ultra-fast computation (below 1ms on a standard desktop and average 3.68ms during onboard flight), while maintaining low onboard memory requirements (less than 1.5MiB). Extensive simulation benchmarks demonstrate superiority in both planning latency and target-reaching progress quality. Zero-shot deployment in real-world experiments further validates the robust sim-to-real transfer capability of the proposed method.

Ruihua Han, Rui Gao, Zhe Liu, ... , Jia Pan, Hengshuang Zhao
总结: 提出形状感知强化学习MPC用于异质形状人群安全导航。
方法: 几何分离特征构建高阶CBF约束,RL学习实时MPC参数更新。
证据: 在密集异质人群中实现安全高效自适应导航无需简化。
为什么适合我: RL-MPC导航相关但非人形腿式接触丰富全身运动。
原摘要

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability. Project website: https://hanruihua.github.io/srl_mpc_project/

Meet Pal Singh, Vyankatesh Ashtekar, Ashish Dutta
总结: 用强化学习从零合成人形机器人自然坐到站运动。
方法: 力与椅高耦合课程、随机IK姿态采样及分阶段奖励驱动PPO。
证据: 策略掌握多高度平滑类人起身轨迹并实现泛化。
为什么适合我: 直接相关人形全身RL控制、课程学习与无参考运动合成。
原摘要

A humanoid has infinitely many ways to stand up from sitting while maintaining balance, making sit-to-stand (STS) a challenging control problem. We synthesise natural humanoid STS motion from scratch using reinforcement learning, without demonstrations or reference trajectories. A single Proximal Policy Optimisation policy learns smooth, human-like rising driven by three complementary components. (i) A coupled force/chair-height curriculum is used. A vertical pelvis-assist force aids early trajectory exploration and decays over training. Taller chairs are unlocked with decaying assisting force. This ensures that the policy masters a viable STS trajectory at each chair height before being exposed to harder ones, avoiding the premature distribution shift that otherwise collapses generalisation. (ii) Motion robustness is achieved by randomly sampling from a large number of inverse kinematics-generated initial and target poses spanning over eight chair heights. (iii) A set of rewards is defined inspired from biomechanics and optimal control studies. They shape the robot's angular momentum for seat-off, and enable support-region transition via centre of pressure attraction function to ensure smooth low-effort actuation. On a deterministic force-free evaluator, the policy attains more than 97% balanced-standing success across eight chair heights. The policy generalises smooth motion across chair heights and enables the robot to rise from substantially deep-seated postures as compared to the state of the art.

Ioannis Papadopoulos, Georgios Tsaousoglou, Johanna Vorwerk
总结: 开发多目标深度强化学习智能体用于电力系统安全稳定。
方法: 统一控制智能体在随机负荷下维持热安全并改善阻尼。
证据: 比单目标智能体更好平衡目标,阻尼改善且热违规可忽略。
为什么适合我: 电力系统多目标控制,与机器人运动控制无关。
原摘要

The ongoing energy transition challenges the stable operation of power systems and increases the need for rapid decision-making under uncertainty. While reinforcement learning has emerged as a promising framework for power system control and operation, existing applications typically focus on a single operational criterion, such as thermal security or small-signal stability. However, power system operation is inherently multi-objective and may involve trade-offs between objectives. This paper develops a unified-control deep reinforcement learning agent that maintains thermal security under stochastic load variations while steering the system toward operating points with improved damping of the most critical mode. Compared to a thermal-security-only agent and a business-as-usual policy, the proposed agent achieves a better balance among the operational objectives considered, with notably improved damping and negligible thermal-security violations. Finally, the operational value of increased critical damping is demonstrated under small- and large-signal disturbances, where operating points with higher damping lead to faster oscillation decay and improved critical clearing times.

Bernhard Hilpert, Kim Baraka, Joost Broekens
总结: 提出TOSS框架建模人机交互学习中人类教学决策过程。
方法: 自下而上研究分析触发目标信号策略网络及教师角色。
证据: 揭示教学为机器人行为与人类动作间程序循环。
为什么适合我: 人机教学框架可启发遥操作,但非全身运动核心。
原摘要

Successful Human-Robot Teaching assumes alignment between robot processing needs and human teaching intent. To better understand this alignment, this work seeks to uncover the underlying logic that humans intuitively apply when teaching. Through an exploratory, bottom-up study with N=34, participants observing two distinct robot Reinforcement Learning (RL) scenarios, we analyze 204 intuitive teaching responses across early, middle, and late learning phases. Results reveal that teaching decisions consist of a nuanced, interconnected network of Triggers (situational catalysts), Objectives (subjective teaching targets), Signals (communicative acts), and Strategies (high-level governance) in which teachers spontaneously adopt diverse roles, acting as coaches, engineers, or designers and prioritize different objectives. Based on these results, we introduce the TOSS Framework, which conceptualizes Human-Robot teaching as a procedural loop between robot behavior and human teaching actions, in which human teaching decisions are modeled as Trigger-Signal responses modulated by teaching Objectives and Strategies. It provides future research with an openly accessible dataset and a theoretical foundation for a) understanding teaching decisions and b) simulating realistic oracles as well as c) designing human-centered teaching settings and novel robot learning algorithms that go beyond the constraints of current robot learning settings.

Varun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev, Animesh Garg
总结: 提出Q规划为BC策略配备离策略Q函数实现机器人自改进。
方法: 用Q估计价值支持推理引导选择,仅微调Q实现在线改进。
证据: 在LIBERO与双臂任务上实现自改进无需额外人类演示。
为什么适合我: 机器人策略自改进与模仿学习可借鉴全身控制器。
原摘要

Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter models underpinning modern robot policies. We propose Q-Planning, which equips a large visuomotor BC policy with a small off-policy Q-function. Because a Q-function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value-guided action selection at inference (a single-step Q-weighted average over BC draws) and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self-improvement lift every benchmark score we tested (LIBERO-10 93% to 99%, RoboTwin 83.8% to 91.4%) and shorten successful episodes on the near-ceiling suites (LIBERO-Object, LIBERO-Goal). On two contact-rich bimanual real-robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack-cups 40% to 90% and insert-wallet 25% to 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q-Planning is the only method, among Best-of-N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor.

Noam Berkovich Lahav, Oren Wiezel, Yizhar Or
总结: 基于Purcell模型优化游泳机器人最优控制与步态。
方法: 校准非细长连杆模型,用庞特里亚金原理求最优周期步态。
证据: 找到关节角约束下最大化每周期位移的最优步态。
为什么适合我: 游泳微机器人最优控制,与腿式人形无关。
原摘要

Purcell's swimmer is a well-known planar model of a swimming microorganism, governed by low Reynolds number hydrodynamics, which is comprised of three rigid links connected by actuated rotary joints. This model has been analyzed as a robotic locomotion system governed by first-order nonlinear dynamics with a periodic input (gait) of the two joint angles. In this work, we present a robotic macro-scale realization of this three-link swimmer moving in a highly viscous fluid. We propose a simple variant of Purcell's theoretical model with non-slender links and a central rigid sphere which represents the added drag of the robot's central flotation block, and calibrate the model's parameters to fit experimental measurements. Next, we apply optimal control formulation based on Pontryagin's Maximum Principle (PMP) in order to find optimal gaits that maximize the displacement per cycle under bounds on the joint angles. Employing a differential geometric method that transforms the problem to area integral enclosed by the gait trajectory in the plane of joint angles, enables visual interpretation which explains topological changes in displacement-optimal gaits upon varying the bound on the joint angles. We then apply PMP formulation to the problem of maximizing Lighthill's energy efficiency in order to obtain a boundary value problem (BVP) whose solution gives efficiency-optimal gaits for Purcell's swimmer model, as well as its variant with a central sphere. Finally, we utilize numerical methods such as parameterizing the input gait as a truncated Fourier series, as well as GPOPS-II solver, to produce sufficient initial guess values for solving the BVPs and obtaining efficiency-optimal gaits.