Jayesh Jangid, Pulkit Suman, Vishal Shrivastava, Mukesh Kumar Mishra, Sangeeta Sharma · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23185914
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision-Language Models (VLMs) represent a significant advancement in Multimodal Artificial Intelligence by integrating Computer Vision and Natural Language Processing into a unified framework. These models are capable of understanding visual and textual information simultaneously, enabling intelligent reasoning, image understanding, visual question answering, image captioning, and human-computer interaction. Recent developments in Large Vision-Language Models (LVLMs) have further enhanced the capability of AI systems to perform complex multimodal tasks with high accuracy and contextual awareness. This review paper presents a comprehensive study of Vision-Language Models, including their architecture, working principles, training methodologies, and recent advancements. The paper examines popular VLMs such as CLIP, BLIP, LLaVA, GPT-4o, Gemini, and Qwen-VL, along with their applications in healthcare, education, robotics, autonomous systems, and assistive technologies. Furthermore, the study discusses the advantages, limitations, and challenges associated with VLMs, including computational complexity, data requirements, bias, and hallucination issues. Finally, future research directions and emerging trends in multimodal AI are highlighted. [1,2]The review concludes that Vision-Language Models are transforming the field of Artificial Intelligence by enabling machines to understand and interact with the world in a more human-like manner. Keywords: Vision-Language Models (VLMs), Large Vision-Language Models (LVLMs), Multimodal Artificial Intelligence, Computer Vision, Natural Language Processing, Deep Learning, Image Understanding, Visual Question Answering, Image Captioning, Human-Computer Interaction, Large Language Models, Multimodal Learning.
No comments yet — start the discussion below.