Zayr Habeeb, Sk Miraj Ahmed · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23041275
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background/Motivation. Multimodal medical AI agents such as MedRAX [1] combine vision-languagemodels with specialist diagnostic tools to answer clinical questions from chest X-rays, yet their failuremodes remain poorly characterized, and it remains unclear whether such failures can be repaired atinference time without retraining. Methods/Approach. We developed a five-stage hierarchical failure taxonomy spanning perception,reasoning, evidence, generation, and safety, and applied it to 1,495 annotated cases across three models—MedRAX, Qwen2.5-VL-7B-Instruct, and CheXagent—on ChestAgentBench [1]. Motivated bylimitations in MedRAX tool use, we designed Ensemble-Vote Repair, which combines fivecomplementary answer signals: the original prediction, three independently resampled predictions, andone prediction reconsidered using evidence from specialist diagnostic tools. The original answer isreplaced only when the ensemble satisfies a specified voting threshold. We evaluated the approach on a495-case sample and the full 2,429-case benchmark and systematically varied the voting threshold tocharacterize the tradeoff between correcting initially wrong answers and introducing errors into initiallycorrect predictions. Results. Diagnostic tool invocation by MedRAX was nearly absent despite the availability of specialisttools. Ensemble-Vote Repair improved accuracy from 44.65% to 46.06% on the 495-case sample andfrom 43.19% to 45.78% on the full benchmark. A systematic voting-threshold sweep showed that a morepermissive majority rule outperformed the initially deployed configuration at both scales. Directly trustingthe tool-informed reconsideration alone corrected approximately 25% of initially wrong predictions butalso changed approximately 41–42% of initially correct predictions to incorrect ones, demonstrating thattool evidence should be corroborated before overriding an existing answer.Significance/Broader Impact. These results show that inference-time reliability methods for multimodalmedical agents must balance error recovery against newly introduced errors. Combining specialist-toolevidence with controlled agreement provides a practical strategy for improving reliability withoutretraining the underlying agent and could extend to other tool-using multimodal AI systems.
No comments yet — start the discussion below.