Prof. Kiran B, Sandesh, Yashwanth L, Yashwanth M, Yashwanth P · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23157872
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Digital information in regional languages is expanding rapidly, but accessing such material becomes difficult when the user's query and the available documents are written in different languages. CrossLinguaRAG is a web-based Cross-Language Information Retrieval (CLIR) system developed to retrieve information from Kannada documents using either English or Kannada queries. The system accepts Kannada PDF, DOCX and TXT files and processes them through text extraction, language validation, cleaning and overlapping chunking. Each passage is represented with the pretrained paraphrase-multilingual-MiniLM-L12-v2 Sentence-BERT model. Retrieval uses a hybrid relevance strategy that combines semantic similarity, TF-IDF lexical similarity and keyword matching. The highest-ranked Kannada passages are returned as evidence and can be translated into English. A Retrieval-Augmented Generation (RAG) stage then uses the retrieved evidence and the original query to produce an English response while keeping the supporting passages visible. A manually annotated evaluation set of 20 English and Kannada queries was measured using Recall@1, Recall@3, Mean Reciprocal Rank (MRR) and nDCG@3. The prototype obtained 80.0% overall Recall@1, 90.0% Recall@3, 0.8333 MRR and 0.8695 nDCG@3.
No comments yet — start the discussion below.