Leonardo Alfredo Forero Mendoza, Harold Dias de Mello, Manoela Kohler, Evelyn Batista, Marco Aurelio Pacheco · Artefactum 2026 · 2026
DOI: 10.23900/artefactum.v25i7.4443
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper presents a multi-benchmark evaluation of the Deep Reinforcement Learning Market-Driven Hierarchical Neuro-Fuzzy Politree (DRL-MD-HNFP) [13,14] against six modern multi-agent reinforcement learning (MARL) baselines [1,4,10,19,20,22], combined with a high-resolution controlled experiment providing intervention evidence for reward-driven performance modulation. Across three benchmarks — Pursuit Game [12], FootballGrid, and RWARE [17] — no single algorithm dominates universally: DRL-MD-HNFP matches the best Q-learning method in shaped-reward environments and is statistically indistinguishable from IQL on the official PettingZoo MPE benchmark [21] (d=0.45, p=0.442, ns), while policy gradient methods [4,22] dominate the sparse-reward sequential RWARE task [17]. A ten-condition controlled experiment yields Spearman ρ=0.964 (p=0.000007), providing robust statistical evidence that reward structure [16] modulates which learning paradigm is viable for coordination. A full-protocol ablation (15 seeds, 1000 episodes) finds no significant performance difference between the fuzzy role layer and random or learned embeddings (Kruskal-Wallis H=3.11, p=0.376), correctly relocating the architecture's primary contribution from performance to interpretability [7]. A Diagnostic Steps to Failure Identification (DSFI) proxy metric is introduced to quantify interpretability gains.
No comments yet — start the discussion below.