Basile Bete Mbezele, Ghislain Alo’o Abessolo · Complex & Intelligent Systems 2026 · 2026
DOI: 10.1007/s40747-026-02494-y
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multi-agent reinforcement learning (MARL) in partially observable and non-stationary environments requires agents to simultaneously infer hidden states, coordinate under limited communication, and learn stable cooperative policies. While existing approaches typically address these challenges through independent mechanisms, their interactions remain insufficiently exploited. We present H3C-BEACON ( Hierarchical Hybrid Heterogeneous Control with Bayesian-Elite Adaptive Coalition Network ), a unified hierarchical framework for coordination and control in cooperative MARL that jointly integrates communication, probabilistic belief inference, adaptive coalition formation, and policy stabilisation. The framework combines six complementary components: (i) Dynamic Graph Attention Networks (DGAT) for distance-aware communication, (ii) Bayesian belief fusion for hidden-state estimation under partial observability, (iii) spectral coalition formation for adaptive role specialisation, (iv) a dual-critic architecture that separates global coordination from local decision making, (v) RTD++ elite-trajectory anchoring to improve policy optimisation stability, and (vi) bounded entropy control to balance exploration and exploitation. We evaluate H3C-BEACON on three widely used cooperative MARL benchmarks spanning communication-intensive, coordination, and imperfect-information settings. On the Multi-Agent Particle Environments, H3C-BEACON consistently improves coordination quality over MAPPO, achieving a perfect win rate across all five independent random seeds on while improving the best episode reward from $$-6.06\pm 0.70$$ - 6.06 ± 0.70 to $$-2.35\pm 0.62$$ - 2.35 ± 0.62 . On , the framework substantially reduces performance variability, producing a $$95\%$$ 95 % confidence interval approximately $$28\times $$ 28 × narrower than MAPPO ( $$\pm 0.57$$ ± 0.57 versus $$\pm 15.90$$ ± 15.90 ), indicating significantly improved reproducibility across random initialisations. Ablation experiments indicate that each architectural component contributes to overall performance, with win-rate reductions ranging from $$28\%$$ 28 % (−DGAT) to $$70\%$$ 70 % (−RTD++ or −Coalitions) when individual modules are removed. Under partial observability in Hanabi-full, H3C-BEACON increases the mean score from $$2.29\pm 0.23$$ 2.29 ± 0.23 to $$3.96\pm 0.82$$ 3.96 ± 0.82 ( $$+73\%$$ + 73 % ) while avoiding policy collapse across all runs, suggesting that RTD++ provides effective stabilisation during cooperative learning. Although MAPPO remains superior on StarCraft combat scenarios, a result consistent with the structural characteristics of that environment (homogeneous units, dense global state, and no explicit communication channel that would benefit from DGAT or coalition formation), the overall results indicate that jointly modelling communication, belief estimation, adaptive coalition formation, and stable optimisation provides a robust and effective framework for cooperative MARL in environments characterised by partial observability and decentralised coordination. All primary results use 5 independent random seeds with $$95\%$$ 95 %
No comments yet — start the discussion below.