
Mallikarjunarao Sunke, Srikanth Gudi, Sriharsha Gudi · International Journal for Research in Applied Science and Engineering Technology 2026 · 2026
DOI: 10.22214/ijraset.2026.84791
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Thanks to the development of basic models and the high quality of the data, the emergence of AI-generated content has accelerated. Despite its incredible success, there are still challenges that are yet to be addressed in AI content generation, such as the processing of long-trail information, maintaining up-to-date knowledge, addressing high inference and training expenses, and addressing data leakage. The shift to address those challenges has been called Retrieval Augmented Generation (RAG). RAG has brought the process of gathering information, which improves the process of data generation by recovering relevant information from available data sources, resulting in robustness and accuracy. RAG has become the foundation of Natural Language Processing (NLP) to effectively fill the gap between factual accuracy of knowledge and fluency of “Large Language Models (LLMs)”. This study traces the root of RAG from its beginnings as a framework for knowledge-based tasks to its current state as an agentic, modular and complex runtime of knowledge. This study is an in-depth analysis of the evolution of Naïve RAG to Modular and Advanced RAG models, and the introduction of new innovations, such as self-reflection, dense vector recovery, and the use of different models. They are then examined to provide detailed feedback on how to make RAG truly dynamic and usable as a verifier when applying them to organizations.
No comments yet — start the discussion below.