Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M. Sutter, Julia E. Vogt, Jonas Kluckert, Thomas Frauenfelder, Christian Blüthgen, Farhad Nooralahzadeh, Michael Krauthammer · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-66181-1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks such as report generation or abnormality detection, they often lack support for interactive diagnostic capabilities. In this work we present RadVLM, a compact, multitask conversational VLM for CXR interpretation. We construct and standardize a large-scale CXR instruction dataset comprising over 1 million image-instruction pairs from multiple public datasets. The dataset integrates single-turn tasks—including report generation, abnormality classification, and visual grounding—with synthetic multi-turn conversations generated from structured CXR attributes. After fine-tuning RadVLM on this instruction dataset, we evaluate it across different tasks together with re-implemented baseline VLMs. Among the evaluated baselines, RadVLM achieves the strongest performance in conversational capabilities and visual grounding, while remaining competitive in other radiology tasks. Ablations comparing task-specific fine-tuning with full multitask fine-tuning are consistent with a benefit of joint training, particularly for lower-resource grounding and conversational settings. Together, these findings support RadVLM as a research prototype for structured CXR interpretation and conversational capabilities to support more effective and accessible diagnostic workflows.
No comments yet — start the discussion below.