Ernesto Moltó, Desireé Ruiz, Vanessa Moscardó, Yudith Cardinale · Big Data and Cognitive Computing 2026 · 2026
DOI: 10.3390/bdcc10100338
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
With the goal of facilitating efficient access to data to support the development of new R&D&I initiatives, this paper presents an information retrieval RAG-based system for projects funded by the Horizon 2020 program domain through multiple configurations and specific adaptations, considering CPU-only hardware demand. This work is an extension of a previous proposal in order to address the limitations when moving from exploratory prototypes to robust, large-scale experimentation. Additionally, we propose an automatic evaluation pipeline that enables large-scale, reproducible comparison of different configurations by analyzing generated answers and automating the scoring process. A methodology structured into three main phases guided the development of the system. First, project embeddings were generated using vector representation techniques, segmenting documents to compare models in terms of semantic quality. Second, relevant fragment retrieval was performed using language models to assess their relevance to a query. Third, system performance was evaluated using language model-based evaluators, automating the process. Additionally, an in-depth analysis identified evaluators that best approximate human judgment, comparing their results with metrics from the reference study. This work contributes to the State of the Art by proposing improvements in information retrieval systems.
No comments yet — start the discussion below.