Chaitra Gopalappa, UMass Amherst, Sonza Singh · IISE Annual Conference & Expo 2026 · 2026
DOI: 10.21872/annual2025_5605
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multi-agent reinforcement learning (MARL) methods are suitable tools for modeling highly dynamic decision analyses such as intervention analyses in epidemic outbreaks across multiple jurisdictions. MARL models allow us to consider jurisdictional mixing interactions in the decision-making policies, as they evaluate decisions specific to each jurisdiction while considering the interactions between all jurisdictions. In multi-agent reinforcement learning, there are three commonly used training and execution paradigms. Centralized training with centralized execution (CTCE) converts a multi-agent problem into a single-agent problem by considering the joint state/action space of all agents as that of a single virtual agent. Centralized training with decentralized execution (CTDE) trains agents using centralized information but executes their actions independently based on their local observations. Decentralized training with decentralized execution (DTDE) trains agents independently to optimize team rewards, and each agent regards other agents as a part of the environment. We propose a new training and execution approach, where a single deep learning model is trained for all agents (which considers the interactions between agents) but each agent makes its own independent decisions. The standard DTDE approach trains a separate network for each agent. This approach is computationally complex, and this complexity grows with the number of agents. The new method, which we call decentralized training with decentralized execution using a shared network (DTDE-SN) models interactions between agents while reducing computational complexity. We have used a simple epidemic simulation model to demonstrate that the DTDE-SN method can identify optimal intervention decisions.
No comments yet — start the discussion below.