Papers for 2026-09-27

10 papers

Variable-Horizon Model Predictive Control for Switched Systems

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Rui Zhao, Zhiqiang Zuo, Yang Shi, ... , Zheng Li, Guanrong Chen
总结: 变时域切换MPC放松驻留时间要求并保证稳定性。
方法: 构建切换可行集解耦约束并设计短时域方案。
证据: 仿真实验验证了方法的有效性。
为什么适合我: 可借鉴切换MPC处理机器人复杂接触场景。
推荐理由: 切换系统的变时域 MPC 理论,与人形/腿足全身控制无关。
原摘要

This paper investigates model predictive control (MPC) for switched systems subject to control and state constraints. A variable-horizon switched MPC approach is proposed. By steering the system state into a well-designed switching feasible set, the proposed method structurally decouples the dwell-time conditions from the MPC constraints, thereby relaxing the dwell-time requirements to match those of the unconstrained switched systems. Furthermore, algorithms are developed to construct this switching feasible set and characterize the domain of attraction, ensuring both persistent feasibility and closed-loop asymptotic stability. To further decouple the prediction horizon length from strict dwell-time bounds, advanced short-horizon switched MPC schemes are designed, which expand the overall domain of attraction. Simulations illustrate the efficacy of the proposed methods.

Accelerating Branch MPC with Two-Level Parallel Direct Solves on GPUs

1.0/5 偏低 裁判分 2.0 Candidate 当日相对
Fenglong Song, Luyao Zhang, Liang Wu, Ján Drgoňa, Colin N. Jones
总结: GPU两级并行直接求解加速分支MPC。
方法: 场景与时域并行Cholesky分解并定制变量排序。
证据: 数值实验显示相对cuDSS分解加速达6.0倍。
为什么适合我: 加速MPC求解支持全身控制真机实时运行。
推荐理由: 分支 MPC 的 GPU 直接法线性求解器,是数值计算而非机器人运动控制。
原摘要

Branch model predictive control optimizes multiple future trajectories coupled through shared decisions, with computational demands increasing as the number of scenarios and prediction horizon grow. We present a GPU-accelerated direct linear solver for branch MPC formulations in which all trajectories share a single root decision node and evolve independently thereafter. By operating at the linear-algebra level, the solver provides a reusable backend for multiple optimization algorithms whose reduced systems have the required symmetric positive-definite structure. The solver exploits two levels of parallelism: across scenarios and along each prediction horizon. A tailored variable ordering enables horizon-parallel Cholesky factorization while preserving a single root-tail coupling block per scenario in the factor. Numerical experiments demonstrate substantial speedups over state-of-the-art sparse direct solvers, achieving factorization speedups of up to 6.0$\times$ over cuDSS and 27.6$\times$ over eight-thread PARDISO, with triangular solve speedups of up to 3.5$\times$ and 15.8$\times$, respectively.

James Wu, Chris R. Sims
总结: MI-SARSA引入互信息正则建模生物有界理性。
方法: 学习边缘先验并惩罚状态特异动作偏差。
证据: 信息成本可预测反应时间并区分标准RL。
为什么适合我: 有界理性策略可启发机器人模仿与强化学习。
推荐理由: 有界理性与策略复杂度的 RL 理论模型,不涉及机器人全身或腿足控制。
原摘要

Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior. This yields a sequential learning model in which state information is used selectively when its expected return benefit justifies the added informational cost. Critically, the same state-specific information cost that governs policy compression also generates trial-level predictions for reaction time, distinguishing MI-SARSA from most reinforcement learning models, which predict choices or returns but not latency. Empirically, MI-SARSA produces a reward-complexity tradeoff, and stronger information penalties produce simpler policies with lower control costs and faster reaction times. Under environment shift, increasing regularization reduces post-switch performance degradation but also lowers asymptotic return, revealing a robustness-capacity tradeoff. Together, these results position MI-SARSA as a model of bounded sequential learning under cognitive constraints.

A Field-Deployable GNSS-based Navigation Stack for Outdoor Mobile Robots

1.0/5 偏低 裁判分 1.0 Candidate 当日相对
Yiyuan Lin, Cole Regnier, Yu Jiang
总结: 野外可部署GNSS导航栈支持多定位与控制器。
方法: ROS2统一接口含纯追踪NMPC及混合调度。
证据: 葡萄园800次实地运行评估八种组合。
为什么适合我: 户外定位导航可参考腿足复杂地形感知。
推荐理由: 户外移动机器人 GNSS 导航栈,偏自主导航,不在腿足感知运动或全身控制兴趣内。
原摘要

Outdoor robots require more than an accurate receiver and a path-tracking law: the navigation system must preserve geometric consistency from geographic waypoints to actuator commands, expose measurement validity and timing, and respond to invalid or stale state information. This work presents a ROS~2 navigation stack with interchangeable single-GNSS--IMU and dual-antenna-GNSS localization front ends. Both provide a common local East--North--Up state interface for pure pursuit, virtual-point cross-track PID, finite-horizon nonlinear model predictive control (NMPC), and a segment-dependent hybrid dispatcher. The architecture specifies coordinate conventions, datum initialization, asynchronous state construction, waypoint geometry, controller equations, quality gates, command arbitration, and watchdog behavior. Independent physical field runs collected during 2025 and 2026 grape-vineyard deployments support a balanced evaluation of 800 runs, with 100 runs for each of eight controller--localization combinations on an approximately 199.6-m route. The row-hybrid mode yields the lowest run-averaged post-acquisition mean absolute cross-track error (MAE) in the evaluated dataset: 0.00952~m with single GNSS+IMU and 0.00846~m with dual GNSS. These findings characterize deviations of the recorded positions from the reference route under the evaluated conditions. The open-source navigation software and deployment instructions are available in the https://github.com/YiyuanLinXX/PPBv2/tree/main/PPBv2_Navigation.

Alan Royce Gabriel Samuel, Pulkit Verma
总结: 混合刚柔臂耦合建模量化解耦损失并蒸馏策略。
方法: 推导含气动迟滞模型比较MPC与RL后蒸馏。
证据: 耦合控制跟踪精度为解耦PID的2.5倍。
为什么适合我: 耦合MPC与蒸馏利于loco-manipulation控制。
推荐理由: 刚柔混合气动折纸机械臂的建模与控制,落入明确不感兴趣的软体机器人。
原摘要

Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of $N$ alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track $2.5\times$ tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94$\%$ of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows' lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.

agentic-ger: terminology recovery in long-form speech using global context

1.0/5 偏低 裁判分 0.0 Candidate 当日相对
Yanqiao Zhu, Wupeng Wang, Zhifu Gao, Xiangang Li, Xie Chen
总结: LLM代理用全局上下文修正长语音术语错误。
方法: 识别可疑词选择性重转录并引导后续编辑。
证据: 中文实验相对Whisper B-CER最高降36.8%。
为什么适合我: 语音识别与机器人全身运动智能无关。
推荐理由: 长语音术语纠错与ASR,和人形全身运动智能无关。
原摘要

Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses. It selectively re-transcribes the source speech to check candidate corrections, and uses accepted edits to guide subsequent decisions. Experiments with four LLMs and two ASR systems on GigaSpeechBench show consistent terminology improvements in both Chinese and English, with and without thinking. On Chinese speech, Agentic-GER achieves up to a 36.8% relative reduction in biased character error rate (B-CER) over the Whisper baseline.

Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao
总结: TOLA实现文本图像超分的一步潜在适应。
方法: 置信加权文本条件与潜在残差校正模块。
证据: 避免错误文本先验反复注入放大误差。
为什么适合我: 扩散一步适应可启发运动生成但领域不同。
推荐理由: 文本图像超分与OCR条件扩散,和人形全身控制、腿足运动或动作重定向无关。
原摘要

Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.

Pu Wang, Yongcong Wang, Wenhao Li, ... , Shujun Fu, Zhuoran Zheng
总结: FluidRain用无散雨流引导循环注意力去雨。
方法: 估计无散流场引导多尺度跨帧注意力。
证据: 方法设计弥合时序聚合与雨运动建模差距。
为什么适合我: 视频去雨与腿足感知运动控制关联较弱。
推荐理由: 视频去雨的视觉方法,不涉及腿足感知运动或全身控制。
原摘要

Existing video deraining methods typically exploit neighboring frames through either explicit alignment or implicit spatiotemporal aggregation. Explicit alignment relies on accurate motion estimation, which can become unreliable under dense rain, while implicit aggregation avoids alignment but lacks explicit guidance on the directional and temporally coherent structure of rain. This leaves a gap between reliable temporal aggregation and explicit modeling of rain motion. To address these limitations, we propose FluidRain, a lightweight video derainer that uses divergence-free rain flow to guide Loop-in-Loop attention across scales and neighboring frames. Motivated by fluid mechanics, we model rain motion as a divergence-free image-space flow and use it to organize multi-scale and temporal aggregation. Specifically, FluidRain first estimates a rain-flow field for each frame and projects it onto the divergence-free subspace. The resulting flow steers window attention along rain streaks, enabling neighboring frames to be aggregated without explicit alignment. Since rain-flow structure is preserved across scales and nearby frames, Loop-in-Loop reuses the same attention operator across both dimensions, resulting in a three-frame model with only 0.80M parameters. Experiments on four benchmarks show that FluidRain remains competitive with substantially larger restoration models. We further examine how temporal evidence scales with different input views. To evaluate whether the model remains reliable when rain motion changes across frames, we introduce RainSyn-Gust, which injects controlled changes in rain-streak direction into existing benchmarks. We also develop a physics-based no-reference metric that evaluates real-rain removal without requiring clean targets.

Yoshinobu Obata, Yinlai Jiang, Hiroshi Yokoi, Shunta Togo
总结: 仿人软体前臂独立腕骨实现自适应刚度调制。
方法: 解剖准确前臂测不同激活与骨骼配置刚度。
证据: 刚度椭圆与人类测量一致验证形态因素。
为什么适合我: 类人刚度调制启发接触丰富全身柔顺控制。
推荐理由: 属于明确排除的软体机器人,研究仿人腕部骨骼与刚度调制,不是腿足感知运动、跑酷或全身loco-manipulation。
原摘要

The human wrist exhibits adaptive stiffness modulability: joint stiffness anisotropy can be actively regulated through muscle co-contraction. This functionality is essential for stable manipulation, yet the underlying morphological factors remain unclear. To identify these factors, we developed an anatomically accurate anthropomimetic soft robotic forearm comprising eight independently movable carpal bones interconnected by ligaments, 22 actuated muscles, and compliant fingertips. We measured wrist joint stiffness under four muscle activation patterns across three skeletal configurations: anatomically normal carpal bones, a fused proximal carpal row, and a geometric ellipsoidal skeleton. The stiffness ellipse exhibited low stiffness along the dart-throwing motion (DTM) direction when finger muscles were activated, but high stiffness along the same direction when wrist and finger muscles were activated simultaneously. These results agree with previously reported human measurements, demonstrating that precise anatomical replication reproduces human-like stiffness modulability. Fusing the proximal carpal row eliminated the low DTM-direction stiffness under finger muscle activation, while the geometric ellipsoidal skeleton showed poor stiffness ellipse reorientation across all conditions. Carpal bone motion analysis revealed significantly opposing coupling patterns between wrist and finger muscles at the proximal carpal row, accompanied by a consistent but non-significant trend at the midcarpal joint, providing a mechanical explanation for this modulation. These findings demonstrate that carpal bone morphology plays a dominant role in human wrist stiffness modulation and provide design principles for humanoid robot wrists.

Tobias Gergs, Rouven Lamprecht, Sahitya Yarragolla, ... , Thomas Mussenbrock, Jan Trieschmann
总结: 多尺度框架关联忆阻器工艺条件与功能行为。
方法: 统计与物理模拟揭示缺陷到功能概率级联。
证据: 五万余器件聚类显示氧空位为潜在描述符。
为什么适合我: 忆阻硬件与人形机器人运动控制无直接关联。
推荐理由: 忆阻器件材料与工艺,和机器人运动控制无关。
原摘要

Resistive switching in oxide-based devices is widely governed by stochastic defect processes, yet a predictive link between fabrication conditions and functional behavior remains elusive. Here, we establish a multiscale framework connecting plasma-defined deposition conditions to macroscopic device functionality in sputtered SiO$_x$/Cu/SiO$_x$-based systems. By combining large-scale statistical analysis of more than 50,000 experimentally characterized devices with physics-based plasma and atomistic simulations, we show that device behavior does not emerge from deterministic process-to-performance mappings, but from a probabilistic cascade spanning defect formation, defect-state evolution, and functional-regime emergence. Data-driven clustering reveals a continuous functional state space composed of operational switching types, while inverse modeling identifies the reconstructed oxygen-vacancy density as an effective latent descriptor capturing the combined influence of structural disorder and defect topology. This latent descriptor is strongly coupled to both Cu redistribution and electrical response, linking otherwise hidden material properties to observable device characteristics. Furthermore, macroscopic switching behavior is argued to arise from ensemble integration across spatially heterogeneous subdomains, providing a physical explanation for the pronounced variability of large-area devices. These findings shift the perspective from deterministic defect engineering toward probabilistic defect-state design and establish a physically grounded framework for understanding and controlling functional variability in such oxide-based systems, such as memristive or resistive-switching devices.