Bowen Qu, Weixin Liu, Matthew Murrow, Matthew Burger, Xingyi Guo, Mihir Sachin Vaidya, Susannah L. Rose, Murat Kantarcıoğlu, Bradley Malin, Zhijun Yin · medRxiv 2026 · 2026
DOI: 10.64898/2026.09.23.26363835
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Mammography interpretation requires tying each finding to supporting evidence and withholding judgment when that evidence is unavailable. Vision-language models are increasingly used for image-based medical question answering, yet existing benchmarks suggest that they struggle to jointly support accurate answers, evidence grounding, and appropriate abstention. We propose CARE-MVLM, built on Qwen2.5-VL-7B-Instruct, to address this gap through separate modules for answer prediction, answer-conditioned evidence generation, and answerability selection. On 7,358 triplets curated from CBIS-DDSM, MIAS, and VinDr-Mammo, CARE-MVLM substantially improves lesion-sensitive selective behavior over same VLM backbone baselines and exhibits markedly stronger lesion-sensitive abstention than GPT-5.6 and Gemini-3.5-flash while attaining the highest joint answer, grounding, and abstention reliability in the comparison.
No comments yet — start the discussion below.