Azhar A. Hadi, K. P. Supreethi · Discover Informatics 2026 · 2026
DOI: 10.1007/s44564-026-00022-1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In recent years, multimodal deep learning has made great progress, enabling artificial intelligence to learn across text, audio, images, video, and other data. Unlike earlier unimodal deep learning models trained on limited data, new multimodal models are trained on large collections of mixed media and achieve strong results. This review covers 13 recent pre-trained models from 2020 to 2025, including CLIP, GPT-4V, GPT-4o, SigLIP, ViT, BLIP, Flamingo, PaLI, LLaVA, Florence-2, PaliGemma-2, Gemma-3, and Llama-4, that help connect different types of information for tasks such as alignment, classification, and detection. Moreover, this paper provides a comparative analysis of recent pre-trained MMDL models, examining their specific applications, architectural designs, datasets utilized, and evaluation metrics used in current research. It highlights 120 articles on Multimodal Deep Learning (MMDL), including the challenges and limitations of each model. Additionally, this review covers 11 domains using recent pre-trained Multimodal Deep Learning (MMDL), including Healthcare, Education, Industrial and Manufacturing, Autonomous Systems and Transportation, and E-commerce, with dozens of applications across these domains. Although these models perform better, key problems remain, like data quality, high computational costs, model alignment, and the need for more efficient designs. This review also looks at future needs, such as making models faster, more robust, better at handling a wider range of data types, and more trustworthy. Our goal is to provide researchers and students with a clear view of current opportunities and capabilities in multimodal deep learning.
No comments yet — start the discussion below.