Rajiv Ranganath, Thrisha R., Vismaya M., Akshitha Katkeri · International Journal of Innovative Science and Research Technology (IJISRT) 2026 · 2026
DOI: 10.38124/ijisrt/26sep517
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
When several self-directed learners share one environment, each must improve its own behavior while the others are also changing theirs. This is the setting studied by multi-agent reinforcement learning (MARL), and the agents involved may be working toward one goal, pulling in opposite directions, or some combination of the two. In this review we first lay out the mathematical objects used to describe such interaction, namely Markov games and their decentralized, partially observed variants, and then group the algorithmic literature into three lineages: agents that learn in isolation, methods that factor a team value into per-agent pieces (VDN, QMIX and their descendants), and actor-critic schemes that train against a critic with global knowledge (MADDPG, COMA, MAPPO). We discuss the obstacles that distinguish MARL from its singleagent counterpart, namely a moving-target learning problem, the difficulty of dividing a shared reward among team members, limited local views, and growth of the joint action space, and explain why training with global information but acting on local information has become the standard design. Uses of MARL in real-time strategy games, robot teams, automated vehicles, wireless resource sharing, and the orchestration of language-model agents are reviewed alongside the test suites the community relies on. We close by pointing to unresolved questions concerning scale, cooperation with unfamiliar partners, and deployment safety.
No comments yet — start the discussion below.