Anupam Dhakal, Prashant Pokharel, Sabin Adhikari · European Journal of Applied Science Engineering and Technology 2026 · 2026
DOI: 10.59324/ejaset.2026.4(5).11
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding large language models (LLMs) in external knowledge, but typical deployments introduce substantial energy, latency, and cost overhead due to expensive retrieval and context-processing pipelines. Recent work in "Green AI" and sustainable machine learning has highlighted the need to treat energy efficiency as a first-class quality attribute alongside accuracy and latency, yet systematic guidance for energy-aware RAG design remains limited. This review synthesizes emerging research on energy-efficient RAG, organizing techniques across four stages of the pipeline: retrieval, reranking, caching, and context management. Building on recent RAG surveys and controlled experiments with production-like systems, we examine empirical evidence showing that adjusting similarity thresholds, reducing embedding dimensionality, applying vector indexing, and using lightweight rerankers can reduce energy consumption by 20–60% in realistic workloads, sometimes without sacrificing accuracy. We also integrate cross-cutting work on token reduction, KV-cache management, semantic caching, green prompt engineering, and phase-level energy analysis for LLM inference. The resulting taxonomy highlights concrete design patterns—adaptive retrieval, redundancy-aware context selection, semantic caching of prompts and responses, and babbling suppression—that jointly reduce compute while preserving grounding fidelity. The paper concludes with practical recommendations and open research questions for building future RAG systems that are not only accurate and robust but also environmentally responsible [1-8].
No comments yet — start the discussion below.