Binghai Su · Discover Applied Sciences 2026 · 2026
DOI: 10.1007/s42452-026-09529-6
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study proposes an opponent-modeling-enhanced multi-agent reinforcement learning framework for policy generation in dynamic games under incomplete information. The framework reuses a recurrently inferred opponent belief in policy input, action-value evaluation, prediction-error feedback, and approximate best-response regularization, thereby forming a coupled decision loop rather than an independent predictor–learner pipeline. In the main partially observable stochastic game, 50 principal agents interact with 50 heterogeneous opponents under masked observations, stochastic disturbances, and policy switching. Controlled comparisons are conducted against generic, opponent-aware, game-theoretic, model-based, and transformer-based methods, including DRON, LOLA, MAPPO, and MAT. The proposed framework obtains a normalized cumulative reward of 0.93±0.03, an opponent-action prediction accuracy of 0.87±0.02, a policy variance of 0.018±0.004, and satisfies the predefined operational-stabilization criterion at episode 780±30. Relative to single-agent reinforcement learning and the non-cooperative game reference, normalized cumulative reward increases by 19.2% and 13.4%, respectively. Ablation results identify the respective contributions of recurrent opponent inference, prediction-error feedback, feature construction, and best-response regularization. Transfer evaluation on MPE2 Simple Tag retains the advantage in normalized return, capture-avoidance rate, and policy variance without transferring the four-mode prediction metric defined only for the self-developed environment. Theoretical analysis establishes a one-step policy-deviation bound and a finite-horizon cumulative drift bound under explicit boundedness assumptions; it does not establish global equilibrium convergence in non-stationary games. The results support improved opponent-aware adaptation under the reported simulation conditions, while independent reproduction and real-world validation remain necessary.
No comments yet — start the discussion below.