Roni Ramon‐Gonen, Haya Engelstein · Computation 2026 · 2026
DOI: 10.3390/computation14090215
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal large language models (MLLMs) and vision–language models (VLMs) have rapidly entered medicine, demonstrating promising performance in clinical reasoning, radiology report generation, and visual question answering (VQA). However, many current multimodal architectures and pretrained visual backbones remain fundamentally rooted in two-dimensional (2D) image processing, even though major clinical imaging modalities, including computed tomography (CT), magnetic resonance imaging (MRI), optical coherence tomography (OCT), and echocardiography, are inherently volumetric or temporal. This narrative review examines the transition from 2D vision–language systems to volumetric multimodal AI, tracing the evolution from 2D and slice- or projection-based approaches through sequential and video-like methods to three-dimensional (3D) vision foundation models and native 3D VLMs/MLLMs. We examine their representational and computational trade-offs, evaluation gaps, and clinically grounded benchmarks. Approaches differ substantially in how they represent and preserve 3D information. Slice- and projection-based methods offer computational efficiency but may discard spatial context, whereas sequential and native volumetric approaches increasingly model relationships across the full imaging study. Recent 3D foundation models and multimodal systems demonstrate the feasibility of reusable volumetric representations and language-enabled 3D image interpretation, but face barriers in computational cost, training-data scale, evaluation methodology, and clinical reliability. Only 53% of Med-Gemini-3D reports were judged clinically acceptable, and natural language processing (NLP) metrics such as BLEU and ROUGE correlate poorly with diagnostic correctness. True 3D multimodal medical intelligence remains in its early stages. Future progress requires efficient volumetric representation strategies, clinically grounded evaluation frameworks, standardized benchmarks, and robust cross-institution validation.
No comments yet — start the discussion below.