Papers for 2026-09-01

10 papers
Kango Yanagida, Kazuki Miyazawa, Takato Horii
总结: 学习人形在被动移动椅上全向坐姿运动控制。
方法: 扩展站立速度跟踪环境加椅模型与坐姿奖励学RL策略。
证据: 随机命令下几乎全程跟踪,最佳坐姿优于站立策略。
为什么适合我: 涉及人形接触丰富全身运动,可借鉴坐姿推进与迁移。
原摘要

Humanoid robots with quasi-direct-drive actuators continuously generate joint torque while standing, whereas seated humans delegate weight support to chairs during desk work. As a first step toward seated loco-manipulation, we study omnidirectional seated locomotion on a passive mobile chair, requiring unfixed pelvis-seat contact and intermittent foot-floor propulsion of the robot-chair system. We extend a standard standing velocity-tracking environment with a passive-chair model, seated-state rewards, critic-only chair observations, and task-tailored contact settings. The policy is learned without motion-imitation rewards; its actor uses only proprioception and velocity commands, without contact sensing or chair states. In random-command evaluation, the policies tracked omnidirectional commands through nearly all 20-s rollouts, and the best seated policies could outperform the Standing policy in velocity tracking. Across four training seeds, a $2^3$ full-factorial comparison of symmetry regularization (SY), foot-slip regularization (FS), and command curriculum (CC) showed that FS reduced CoT but increased tracking error and that some FS-only policies converged to stationary local optima. Combining FS with either SY or CC avoided this failure without retuning FS, while SY improved bilateral leg symmetry during longitudinal motion. Direction-resolved analysis showed CoT ordered backward $<$ lateral $\ll$ forward, with planted-leg extension in backward and lateral motion and knee flexion following heel contact in forward motion. The learned policy achieved zero-shot sim-to-real transfer to a Unitree G1 and generated omnidirectional seated locomotion.

Yan Pan, Lingfan Bao, Tianhu Peng, Chengxu Zhou
总结: 实时参数化生成人形机器人情感化全身运动。
方法: 从姿态扩张与能量闭式算V-A,先验组合去噪生成。
证据: 运动跟踪全范围V-A命令,动作情感可实时编辑。
为什么适合我: 人形运动生成与跟踪多样动作相关,可延伸生成式策略。
原摘要

People read a humanoid robot's motion in social settings not only for the action performed but for the affect conveyed. Motion carrying that affect has so far been generated for human avatars, where style is taken from a reference clip or an emotion word, neither of which can be quantitatively parameterized. We present PAMoR, which turns affect into a measured control parameter: a valence-arousal (V-A) coordinate computed natively on robot kinematics. It is obtained in closed form from postural expansion and movement energy, and these measurements serve directly as generation conditions, with no human annotation. An action prior and two affect priors, trained in a shared latent space, are composed at each denoising step: the action prior fixes what is performed, the affect priors modulate how. Whole-body motion rolls out autoregressively on a 29-DoF Unitree G1 in real time, with action and affect both editable. Generated motion tracks the commanded V-A over its full range while text-to-motion fidelity still matches text-only baselines. In a perceptual study, raters identify the commanded emotion on 0.38 of trials, above both baselines and approaching the 0.44 reported for acted human bodies.

Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, ... , Lorenzo Cavallaro, Fabio Pierazzi
总结: 深度强化学习逃避并强化恶意软件检测器。
方法: 标签仅黑盒下学可迁移的样本修改与查询策略。
证据: 七检测器平均成功率78.8%,相对提升20.9%至39.2%。
为什么适合我: 与腿式人形机器人全身运动控制兴趣无关。
原摘要

To determine the real-world effectiveness of machine learning based malware detection, it is vital to evaluate its robustness against highly capable adversaries. However, state-of-the-art attacks do not effectively model realistic adversaries, as they often assume access to privileged information such as the training data, feature space, or confidence scores of the target. In this work, we present Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model. Replicant learns a reusable policy on how to modify a malware sample and when to query the target, which transfers across samples, detectors, and feature spaces. Across seven Android malware detectors and three feature spaces, Replicant is the strongest and most query-efficient approach achieving a mean attack success rate of 78.8%, a relative improvement of 20.9%-39.2% over the state-of-the-art. Furthermore, when used for adversarial training, Replicant also outperforms the state-of-the art by producing detectors with more generalizable robustness. With Replicant we demonstrate that learning the task of evasion not only results in stronger attack performance but, crucially, provides a better signal for hardening malware detectors.

Nan Wang, Mohit Yadav, Jonathan Wulff, ... , Xu Dong, Yiwei Tao
总结: 开源仿真就绪腱驱动灵巧手支持操作学习。
方法: 提供电缆传动仿真模型、驱动映射与RL训练包。
证据: 再现传动与拇指耦合,便于从仿真学灵巧操作。
为什么适合我: 灵巧操作与移动操作相关,可借鉴仿真迁移方法。
原摘要

Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the requirement that a motor fit inside the joint it drives, so smaller and cheaper motors suffice, and one motor can drive several joints through a single cable, so fewer motors are needed. They are also harder to learn on than a direct-drive hand. The underactuated transmission that produces the saving is itself difficult to represent in a simulator, and the joints one cable drives are not independently commandable. We present Aero Hand Open, a tendon-driven anthropomorphic hand that is released simulation-ready. Three things ship with it. A simulation model reproduces the cable transmission itself. An identified actuation map connects that model to the motor commands in both directions, including the three-way coupling of the thumb. A reinforcement learning package trains policies for the hand. Together they let a policy be trained entirely in simulation and run on the hand with no fine-tuning and no state estimation. We release the mechanical design, the simulation model, the identified mapping, the training environment and the deployment stack.

Simone Tolomei, Mayank Mittal, Franco Angelini, ... , Paolo Salaris, Marco Hutter
总结: 多批评家RL实现接触引导非抓取移动操作。
方法: 探索批评家密集接触奖励引导,衰减至任务最优策略。
证据: 评估推箱运椅开洗碗机,真机验证椅子运输。
为什么适合我: 高度契合接触丰富非结构化移动操作与全身控制。
原摘要

Non-prehensile manipulation offers versatile skills for moving and rearranging heavy or bulky objects, particularly when combined with a mobile manipulation platform. However, both model-based and model-free approaches struggle with the complex hybrid dynamics and the sparsity of the contact in these tasks. To address these challenges, we propose a contact-guided exploration strategy implemented within a Multi-Critic Reinforcement Learning (RL) framework. A dedicated exploration critic is trained with a dense contact-seeking reward that guides the end-effector toward meaningful contact points; its influence is progressively decayed to recover a task-optimal policy. We obtain candidate interaction points from a general-purpose grasping algorithm, enabling the exploration mechanism to generalise across various object geometries. We evaluate the approach on multiple tasks, including box pushing, chair transportation, and a dishwasher opening task. Finally, we validate the chair transportation policy through extensive experiments on a quadrupedal mobile manipulator, demonstrating deployable non-prehensile manipulation in the real world.

Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang, ... , Wangbo Zhao, Yang You
总结: 智能体游戏开发作可验证数据引擎扩展世界模型。
方法: 游戏引擎提供碰撞物理可玩性奖励支持RL后训练。
证据: 论证可提供比模糊代理更好的空间生成奖励信号。
为什么适合我: 仿真数据引擎可辅助物理角色运动模仿与策略学习。
原摘要

A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.

Massimiliano Bertoni, Alberto Piccina, Gianni Lunardi, ... , Angelo Cenedese, Giulia Michieletto
总结: 生存核MPC为全驱动多旋翼生成安全最优飞行轨迹。
方法: 结合生存理论与数据驱动,动态包围盒强制避障。
证据: 倾斜六旋翼仿真成功导航杂乱环境且实时计算。
为什么适合我: 属飞行器控制,与腿式人形全身运动兴趣无关。
原摘要

Industrial aerial robotics demands safety guarantees for navigation in unstructured environments while optimizing performance and computational efficiency. This paper presents a method for generating safe pose trajectories for fully actuated multirotors within a Model Predictive Control (MPC) framework, leveraging both viability theory and data-driven methods. Obstacle avoidance is enforced through dynamically computed axis-aligned bounding boxes, providing formal safety guarantees without exhaustive offline reachability analysis. Numerical simulations on a fully actuated tilted hexarotor validate the approach, demonstrating successful navigation in cluttered environments with real-time computational performance.

Xianyi Wu
总结: 论证蒙特卡洛树搜索即每访问蒙特卡洛控制。
方法: 从轨迹采样与动作值更新层面统一两者表述。
证据: 四阶段MCTS可归为策略下采样与每访问更新。
为什么适合我: RL基础方法可启发全身控制器策略学习设计。
原摘要

Monte Carlo Tree Search (MCTS) and every-visit Monte Carlo (MC) control are usually presented as different methods. MCTS is described in the language of search (selection, expansion, simulation, and backup), whereas MC control is described in the language of reinforcement learning (trajectory sampling, return estimation, action-value updating, and policy improvement). This note argues that, at the level of trajectory generation and action-value updating, the distinction is largely terminological. The tree policy and rollout policy can be viewed as the learned and not-yet-learned parts of a single evolving policy; expansion corresponds to first visit and initialization; and backup is the ordinary every-visit Monte Carlo update. Under this interpretation, the four stages of MCTS reduce to two basic operations: trajectory sampling under the current policy and every-visit Monte Carlo updating. In this sense, MCTS is simply every-visit Monte Carlo control expressed in the language and data structure of search. The purpose of this note is expository: to make this equivalence explicit and easier to recognize.

Zihan Wang, Bai Huang, Yang Guan, ... , Naizheng Wang, Shengbo Eben Li
总结: 强化学习混合规划器用于不规则环境任意位姿停车。
方法: 目标相对顶点编码环境,耦合学习弧策略与终端集成。
证据: 在阶乘与长程基准上评估成功与轨迹质量。
为什么适合我: 车辆停车规划,与人形腿式运动控制无关。
原摘要

Automated parking commonly assumes marked slots and short approach maneuvers. Delivery and service vehicles, however, may need to reach an operator-specified pose in an irregular bounded environment from a distant start. Existing learning-based parking planners often rely on local observations, which can restrict long-range route reasoning. To address this problem, we present NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking. NeuralParker encodes full-environment obstacle and boundary geometry in a target-relative vertex representation, allowing the policy to retain route-defining context throughout the approach. It further couples a learned curvature--length arc policy with an in-loop terminal ensemble that selects from diverse cubic Hermite connections using a curvature-regularized cost. We also establish factorial and long-range route-choice benchmarks to evaluate planning success and trajectory quality. Experiments on these benchmarks show that NeuralParker achieves higher planning success and better overall trajectory quality than the evaluated baselines, while ablation studies support the benefits of the target-relative global representation and terminal ensemble. Finally, a real-vehicle evaluation confirms that the planner transfers effectively to real delivery-vehicle perception at a working parking site, planning successfully at low computational cost.

Tong Zhang, Yanfei Su, Shuai Wang, ... , Chengzhong Xu, Huseyin Arslan
总结: 多智能体RL实现低空流体天线网络快速重构。
方法: 电磁数字孪生辅助MARL加两阶段迁移学习。
证据: 案例研究显示框架有效缩小仿真到现实差距。
为什么适合我: 通信网络优化,与机器人全身运动控制无关。
原摘要

Low-altitude wireless networks (LAWNs) integrate terrestrial and aerial platforms to provide ubiquitous communication, sensing, and localization services for unmanned aerial vehicles (UAVs) and electric vertical takeoff and landing (eVTOL) aircraft. However, dynamic air-ground and air-air channels, abrupt blockages, and heterogeneous interference hinder the realization of this goal. Nevertheless, fluid antenna (FA), a cutting-edge multiple-input multiple-output (MIMO) technique, overcomes these challenges by reconfiguring antenna positions to unlock additional spatial degrees-of-freedom. In this paper, towards bringing low-altitude FA networks into reality, we study the fast and high-performance FA reconfiguration for low-altitude FA networks with multi-agent reinforcement learning (MARL). Specifically, we present an electromagnetic digital twin (EM-DT)-assisted MARL framework. To fill the sim-to-real gap, we introduce a two-stage transfer learning framework. Our case study shows that joint FA positions and beamforming optimization can enhance the system sum-rate by 118.5%, compared to the fixed position baseline. This gain comes from the dynamic millisecond timescale reconfiguration of FA arrays and the adaptive steering of beams toward aerial users with mobility.