Wang Liu, Xiaowei Xu, Mingxing Deng, Qinghua Qi · Ocean Engineering 2026 · 2026
DOI: 10.1016/j.oceaneng.2026.128026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multi-USV cooperative search and target interception holds value in maritime security, yet existing methods face three challenges: multi-timescale decision coupling, insufficient heterogeneous observation fusion, and inefficient exploration under sparse rewards. This paper proposes a multi-USV decision-making framework based on Hierarchical Spatio-Temporal Proximal Policy Optimization (HST-PPO) with three contributions. A three-layer spatio-temporal decision architecture spanning strategic, tactical, and execution layers is coupled via soft goal passing and trained end-to-end. A Hierarchical Causal Observation Fusion Network (HCOF-Net) uses causal mask attention and temporal decay encoding to adaptively fuse four categories of observations. A Distributed Random Network Distillation exploration mechanism (DRND-NovelD) combines ensemble novelty estimation with a multi-agent diversity reward to accelerate convergence under sparse rewards. HST-PPO is validated across Python simulation, Unity high-fidelity simulation, and real-world lake trials, and benchmarked against ten baselines spanning six non-hierarchical MARL methods and four hierarchical RL methods adapted to the same setting. HST-PPO achieves a Search Coverage Rate of 82.3% and Interception Success Rate of 80.0%, surpassing the strongest baseline by 1.7 and 1.1 percentage points and outperforming every hierarchical baseline. In real-world trials, HST-PPO retains 96.2% and 81.3% of its simulated Search Coverage Rate and Interception Success Rate, validating its zero-shot sim-to-real transfer capability.
No comments yet — start the discussion below.