
Piyush Sharma, Varsha Sahni, Rijwan Khan, Himanshu Gupta · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-72982-1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal Large Language Models (MLLMs) have become a leading approach to visual reasoning, yet the field still lacks a quantitative synthesis that rigorously identifies the drivers of their performance. In this paper, we evaluate 21 models across 7 standardised benchmarks, yielding 102 model–benchmark observations, using a two-level mixed-effects design in which benchmark scores are nested within models. The multilevel model reveals substantial between-model variance (intra-class correlation coefficient, ICC = 0.801) and a significant positive association between model scale and standardised performance (β = 0.73, p = 0.016). Training strategy also matters: relative to instruction tuning, the pretraining + alignment approach is associated with substantially lower performance (β = −1.88, p < 0.001), while early descriptive trends suggest that the Reinforcement Learning from Human Feedback (RLHF) group may be associated with positive effects, although only two models fall within this group, which prevents any definitive conclusion. Correlations between benchmarks are generally strong, ranging from ρ = 0.63 to 0.95 among the reliably estimated pairs. As a sensitivity analysis, we also report a DerSimonian–Laird aggregate meta-analysis, which yields a near-zero pooled effect (d = − 0.046) and extreme heterogeneity (I² = 96.3%). Robustness checks — Hedges’ g validation, leave-one-out sensitivity, bootstrap confidence intervals, SE-floor sensitivity, and Pareto analysis — support the primary conclusions and indicate that training methodology and model scale jointly shape robust multimodal visual reasoning performance.
No comments yet — start the discussion below.