Milad Rahmati, Nasrin Rahmati · Journal of Electrical Systems and Information Technology 2026 · 2026
DOI: 10.1186/s43067-026-00407-0
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In recent years, the proliferation of autonomous systems at the edge—including robotic swarms, autonomous vehicles, and intelligent sensor networks—has highlighted the urgent need for resilient multi-agent planning under uncertain, decentralized conditions. Traditional multi-agent planning architectures typically assume stable communication links and complete observability, conditions that are rarely met in real-world deployments marked by dynamic network topologies, delayed information exchange, and partial observability. This paper introduces Dynamic Federated Multi-Agent Planning (DFMAP), a novel framework that integrates federated reinforcement learning (FRL) with partially observable Markov decision processes (POMDPs) for real-time decision-making in volatile edge environments. The proposed architecture enables agents to learn locally, share model updates without centralized coordination, and collaboratively build global planning policies even under intermittent connectivity. A probabilistic graph-based model tracks dynamic communication topology shifts, while an asynchronous aggregation strategy with staleness correction ensures robustness to network instability. Experiments conducted in both simulated and semi-realistic edge environments—covering decentralized robotic navigation and collaborative search—demonstrate DFMAP’s consistent superiority over centralized and naively decentralized baselines. Specifically, DFMAP achieves a task success rate of 92.5%, a cumulative reward of 121.8 ± 5.4 by training round 100, and a per-round communication overhead of only 15.2 MB, representing more than a 66% reduction relative to the centralized baseline. Under severe communication stress with 50% link dropout, reward degradation is held to just 11.0%, substantially below the 22.4% and 30.0% observed for competing methods.
No comments yet — start the discussion below.