Papers for 2026-08-23

10 papers

Video2DoorTraversal: Push Door Traversal via Simulated Door Twins

4.0/5 相关 裁判分 8.0 Strong 当日相对
Xincheng Tang, Yiji Chen, Youhan Xie, ... , Xibin Song, Ruigang Yang
总结: 单视频实到仿到真实现轮腿机器人开门穿越。
方法: 重建门孪生,仿真迭代生成演示,训双深度策略。
证据: 五扇真门平均成功率96.57%,零样本80.95%。
为什么适合我: 高度相关loco-manipulation与轮腿全身控制。
推荐理由: 最接近「全身 loco-manipulation 与场景交互」:轮腿移动操作的推门穿越、real-to-sim-to-real 与双深度基座-手臂协调控制,虽非双足人形全身跟踪,仍高度相关。
原摘要

Door opening and traversal is a long-horizon loco-manipulation task that requires precise handle interaction and coordinated base-arm control. We present Video2DoorTraversal, a single-video real-to-sim-to-real framework for wheel-legged mobile manipulators. Given one RGB video of a real door, DoorTwin reconstructs an instance-aligned, articulated, and simulation-ready door twin with realistic geometry and appearance. A simulation-in-the-loop agent converts the recovered articulation into a parameterized skill program and iteratively refines failed rollouts to generate physically executable demonstrations. These demonstrations are used to train ArticuACT, a dual-depth policy that predicts coordinated base, arm, and gripper commands using robot-centric camera conditioning and interaction-aware supervision. With all perception and policy inference running onboard, the system achieves a 96.57% average success rate across five real doors and an 80.95% zero-shot success rate on structurally similar unseen doors, while completing the full approach, opening, and traversal sequence in approximately 13s on average. Project Page: https://video2doortraversal.github.io/.

Cong-Thanh Vu, Yen-Chen Liu
总结: DRL使移动机器人动态调整人类伴随位置。
方法: 交互空间作状态,MPPI结合CBF安全控制。
证据: 摘要未提供具体定量实验结果。
为什么适合我: 涉及移动机器人RL,但非腿足全身运动。
推荐理由: 移动机器人伴随行人的深度强化学习位置自适应,只沾边人机交互导航,不是腿足/人形全身运动或 loco-manipulation。
原摘要

In the field of Human-Robot Interaction (HRI), achieving flexibility in human-accompanying within real-world environments holds great potential for various applications but also poses significant challenges. Traditional methods typically restrict robots to fixed positions relative to humans, such as tracking from behind, in front, or side-by-side, which limits robot adaptability in dynamic workspaces. This study introduces a novel human-companioning strategy that uses Reinforcement Learning (DRL) to enable mobile robots to dynamically adjust their tracking positions according to varying conditions. An interaction space is defined to capture the relationship between the human and the robot while considering the environment, which serves as the basis for state spaces in DRL to assist the robot in adapting to environmental changes. A human-robot companion controller is developed by integrating Model Predictive Path Integral (MPPI) control with Control Barrier Functions (CBF), ensuring that the robot accurately follows the target's movement in both position and orientation while avoiding obstacles and enhancing social acceptance and safety. The proposed approach is evaluated in real-world scenarios, both indoors and outdoors, and compared with other studies. The results show that the proposed method improves the success rate and tracking accuracy by at least 24% and 47%, respectively, while enhancing human comfort. Experiments demonstrate the robot's ability to flexibly accompany a person walking at speeds of up to 1.7 m/s, dynamically adjusting its strategy without being confined to a fixed position. Additionally, the robot respects the human's intimate space to ensure safety, comfort, and effective obstacle avoidance.

Honglie Wang, Jia Sun, Zijun Li, ... , Tingting Gao, Yan-Ming Zhang
总结: 后训练框架提升海报文本编辑保真与布局。
方法: 监督微调结合操作特定奖励优化插入替换。
证据: 摘要描述奖励设计但无定量结果。
为什么适合我: 图像文本编辑与机器人运动智能无关。
推荐理由: 商品海报文字编辑与字形渲染,与人形/腿足运动控制无关。
原摘要

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs. We introduce \textbf{TextRefine}, a task-aligned post-training framework that combines supervised fine-tuning with operation-specific reward optimization to address these complementary failure modes. For text insertion, our text-span-level reward jointly assesses semantic fidelity and target-span coverage, penalizes spatial conflicts with products and existing text, and employs a gated structural constraint to preserve non-text regions. For text replacement, our glyph-level reward leverages the connectionist temporal classification (CTC) posterior of the target character to provide graded supervision for fine-grained defects, including missing strokes, structural deformations, and confusion among visually similar characters. We further introduce \textbf{OpenTextEdit}, a dataset comprising 100K images for text editing in product posters, with multi-text layouts, detailed text attributes, product masks, and challenging low-frequency characters. Extensive experiments on both insertion and replacement demonstrate that TextRefine consistently outperforms the evaluated image editing baselines in textual fidelity, placement reliability, and glyph quality while better preserving source-image content.

Dong Qiang, Tian Yuan, Song Yang, ... , Cheng Cheng, Min Yu
总结: 磁自密封MRF触觉执行器实现高保真扭矩。
方法: 磁静仿真设计,PWM激励与模型基扭矩渲染。
证据: 最大输出600 N·mm/A,10kHz降低滞后。
为什么适合我: 触觉硬件与全身运动控制无关。
推荐理由: 磁流变触觉执行器与力矩渲染硬件,不属于全身运动智能。
原摘要

Accurate and stable torque rendering is essential for safe and perceptive human--machine interaction. Magnetorheological fluid (MRF)-based actuators offer a compact and rapidly controllable solution for haptic feedback, but their practical implementation requires reliable fluid sealing, low-hysteresis excitation, accurate torque control, and stable long-duration operation. This article presents an integrated MRF haptic system featuring a compact magnetically self-sealed rotary actuator, low-hysteresis PWM operation, high-fidelity model-based torque rendering, and stable performance during long-time operation. Magnetostatic simulation guides the arrangement of magnetic and nonmagnetic materials to focus flux in the multidisk torque and permanent-magnet sealing regions, enabling a maximum 600 N$\cdot$mm/A output. Experiments show that higher PWM frequencies reduce hysteresis and improve repeatability. At 10 kHz, the response is represented by a nonlinear model that varies with the direction and speed of torque change. The real-time controller combines feedforward, hysteresis compensation, PI feedback, and sliding-mode correction. Compared with PID, it reduces square-wave overshoot, undershoot, and steady-state RMSE by 77.4\%, 61.9\%, and 68.3\%, respectively. It tracks sinusoidal and biomechanics-model-based references, and a 1.5-h test shows only a 2.5 $^\circ$C rise near the coil with no clear tracking loss. This high-fidelity torque rendering will fundamentally transform human--robot collaboration by making interactions safer, more efficient, and more intuitive.

Illya Havrylov
总结: 资源高效混合RAG处理乌克兰多域文档。
方法: BM25、BGE-M3与交叉编码器重排加量化LLM。
证据: OCR耗时5-7小时,限推理约两小时。
为什么适合我: NLP文档任务与机器人运动无关。
推荐理由: 乌克兰多领域文档的 RAG 问答,与机器人运动无关。
原摘要

This paper describes the system submitted to the UNLP 2026 Shared Task on Multi-Domain Document Understanding. The challenge required extracting precise answers, document IDs, and page numbers from a diverse corpus of Ukrainian PDF documents within a strict 9-hour offline Kaggle execution limit. During evaluation on the hidden private test set, optical character recognition (OCR) of scanned documents emerged as a severe bottleneck, consuming 5-7 hours of the total time budget due to sequential single-threaded execution. This overhead strictly limited the remaining time for Large Language Model (LLM) inference to approximately two hours for 500 questions. To guarantee pipeline completion without timeouts, we developed a resource-efficient Hybrid Retrieval-Augmented Generation (RAG) pipeline utilizing BM25, BGE-M3, and Cross-Encoder reranking. Rather than deploying parameter-heavy reasoning models (e.g., DeepSeek R1) which consistently timed out, we utilized a 4-bit quantized LapaLLM 12B model via llama.cpp on dual NVIDIA T4 GPUs. Prioritizing pipeline stability over multi-step reasoning, our system achieved a Private Score of 0.8095, placing 10th out of 15 active teams.

Poomphob Suwannapichat, Boonyarit Changaival, Caesar Wu, Pascal Bouvry
总结: 奖励引导自回归图生成优化多智能体拓扑。
方法: 训奖励模型捕获正确性与紧凑性后微调。
证据: 保持准确率同时平均减少token消耗20.5%。
为什么适合我: LLM多智能体拓扑与机器人无关。
推荐理由: LLM 多智能体通信拓扑的图生成,不是物理角色或腿足控制。
原摘要

LLM-based Multi-Agent Systems (MAS) achieve strong performance on complex reasoning tasks by coordinating multiple agents, but at the cost of substantial token consumption. Recent work on automatic topology design, ARG-Designer, has reframed this problem as autoregressive graph generation. However, its training objective provides no explicit incentive for the model to generate sparse and efficient topologies. We address this limitation by introducing a Reward-Guided Autoregressive Graph Generation (RGA-Designer) inspired by Reinforcement Learning from Human Feedback (RLHF). We train a reward model that jointly captures task correctness and structural compactness, and then fine-tune the pretrained graph generator using the reward model as feedback. Our method preserves task accuracy at the level of ARG-Designer while reducing token consumption by an average of 20.5%.

End-to-end Early Classification of Time Series in Non-Stationary Environments

1.0/5 偏低 裁判分 0.0 Candidate 当日相对
Aurélien Renault, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire
总结: 端到端RL应对非平稳时间序列早分类。
方法: DQeND统一学习表示、分类与触发决策。
证据: 多漂移场景下稳健优于可分离基线。
为什么适合我: 时间序列RL与腿足机器人无关。
推荐理由: 非平稳时间序列早分类,与运动控制无关。
原摘要

Early Classification of Time Series (ECTS) requires making accurate decisions as early as possible in inherently online and evolving environments. Yet, most existing methods assume stationarity and rely on separable designs, where classification and triggering are optimized independently, an assumption that fundamentally limits their adaptability under drift. In this work, we challenge this paradigm and study ECTS under non-stationary conditions. We provide the first systematic comparison between separable and end-to-end approaches across controlled drifting scenarios. Building on Reinforcement Learning, we introduce DQeND, a unified architecture that jointly learns representation, classification, and triggering decisions, while remaining directly comparable to state-of-the-art separable baselines. Across a wide range of drifts, DQeND demonstrates strong robustness across various non-stationary scenarios, consistently outperforming separable baselines. An ablation study further highlights that jointly updating representation and decision modules is critical to these gains. Overall, our results indicate that end-to-end learning can offer improved adaptation capabilities for ECTS in dynamic environments, and motivate further investigation of alternatives to separable designs.

Bo Qian, Yuting Wu, Shuang Zeng, ... , Dalin Zhang, Jiqiang Liu
总结: 里程碑推断优化长程LLM智能体策略。
方法: 发现里程碑,可靠性校准与进度对比校准。
证据: 无需额外标注,从分组rollout推导信用。
为什么适合我: LLM智能体信用分配与机器人无关。
推荐理由: 长时程 LLM agent 的图策略信用分配,不是机器人全身控制。
原摘要

Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping or graph-based advantage estimation, but can overlook meaningful intermediate milestones. We propose MileGPO (Milestone Inference with Local Evidence for Graph-Based Policy Optimization), which derives process-level credit from grouped on-policy rollouts through three designs. Milestone Discovery identifies candidate milestones on successful rollouts and recurring traps on failed ones. Reliability-Calibrated Shaping (RCS) weights these candidates by outcome-based confidence, strengthening reliable milestones and traps while down-weighting uncertain ones. Progress-Contrastive Calibration (PCC) further tests whether a candidate reflects local progress and whether its incoming transition outperforms observed alternatives from the same state. MileGPO requires neither auxiliary models nor additional environment interaction. Experiments on ALFWorld and WebShop show state-of-the-art performance and a small in-distribution to out-of-distribution gap on ALFWorld. Ablations and credit diagnostics indicate that reliability weighting, local progress, and same-state branch evidence complement milestone discovery and resolve ambiguous intermediate credit.

MidTool: Mid-training Data Synthesis for Agentic Tool Use

1.0/5 偏低 裁判分 0.0 Candidate 当日相对
Fengqing Jiang, Yite Wang, Boyi Liu, ... , Radha Poovendran, Yuxiong He
总结: MidTool构建语料强化LLM工具使用能力。
方法: 结合网页PDF代码与真实工具API合成监督。
证据: 中训后SFT与RL优于基线。
为什么适合我: LLM工具使用与机器人运动无关。
推荐理由: 面向工具调用的 LLM 中训练数据合成,与腿足/人形运动无关。
原摘要

Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as math and science, and can also improve agentic capabilities in software-engineering settings. In this work, we study the parallel but less explored agentic capability: general tool use. We present MidTool, an open corpus construction pipeline for agentic tool-use mid-training that combines large-scale web, PDF, and code data with synthesized supervision from real-world tool APIs, MCP skills, and document-grounded workflows. MidTool is designed to teach models how to recognize tool affordances, ground arguments from context, compose tool call workflow, and recover from incomplete information. We mid-train Qwen3-4B-Base and Qwen3-8B-Base on MidTool-Mix, and then apply follow-up post-training with both supervised fine-tuning and reinforcement learning. Compared with baselines, MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe. These results suggest that general tool use, like other important LLM capabilities, benefits from dedicated mid-training rather than being left entirely to post-training.

Bashirul Azam Biswas, Amartya Bhattacharya, Biratal Raj Wagle, ... , James B. Yu, Indrani Bhattacharya
总结: 多模态自监督跨示踪剂实现PET/CT病灶分割。
方法: 上下文感知掩码重建,预训练后微调分割。
证据: 多机构多癌多示踪剂数据验证泛化。
为什么适合我: 医学影像分割与机器人无关。
推荐理由: 全身 PET/CT 病灶分割的自监督学习,属医学影像。
原摘要

Deep learning-based whole-body PET-CT lesion segmentation can support cancer staging, treatment planning, and response assessment, but generalization is limited by scarce annotations and domain shifts. Self-supervised learning (SSL) can address these challenges but remains underexplored in pan-cancer, multi-tracer PET-CT. In this work, we propose MUST-PET (MUltimodal Self-Supervised learning across Tracers), a multimodal, multi-tracer SSL framework for generalizable whole-body PET-CT lesion segmentation. MUST-PET is trained and validated on a diverse, multi-institutional collection of pan-cancer PET-CT scans acquired with FDG and prostate-specific membrane antigen (PSMA)-targeted radiotracers. MUST-PET uses context-aware masked reconstruction, where one modality is partially masked and reconstructed using complementary information from both PET and CT. The pretrained model is subsequently fine-tuned with labeled samples and evaluated for reconstruction quality, lesion segmentation, label efficiency, and generalizability across independent held-out datasets. MUST-PET reduces reconstruction error, improves lesion segmentation over training from scratch, and performs well with limited labeled data and on unseen external datasets, demonstrating the potential of multi-tracer SSL for label-efficient, generalizable whole-body PET-CT. segmentation.