
Ziming Gao, Weiwei Lu, Wanting Zhu, Wenke Xia, Ruixue Tian, Weiqi Li, Peiming Zhang · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-72525-8
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Although large language models have shown great promise in the medical domain, they still face challenges in complex medical reasoning tasks, including hallucinations and inconsistent reasoning. To address these challenges, we propose MRER (multi-agent reasoning with evidence retrieval), a multi-agent retrieval and reasoning framework inspired by evidence-based medicine. MRER employs a closed-loop adaptive reasoning process that uses accumulated evidence to identify unresolved evidence needs and guide targeted follow-up retrieval. Experimental results show that MRER achieves an average accuracy of 70.68% across three widely used medical benchmarks, representing an absolute improvement of 8.20% over the direct inference baseline. The framework enables a lightweight 8B open-source model to outperform a 70B medical domain-specific model and GPT-3.5 in our experiments. MRER also demonstrates adaptive, on-demand computation when handling complex queries. Human evaluation further indicates that MRER improves the logical coherence and trustworthiness of the generated reasoning. These findings suggest that incorporating evidence-based medicine principles into multi-agent collaboration can support the development of more evidence-grounded and scalable medical AI systems.
No comments yet — start the discussion below.