Mohamed Ben Salha, Fiete Lüer, Maik Betka, Stefan Wagner · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.37749
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The increasing importance of Information Retrieval (IR) in managing large datasets has highlighted significant limitations in traditional keyword-based search systems. Context-aware chat-based search methods, such as Retrieval Augmented Generation (RAG), have recently emerged, but their evaluation compared to keyword-based systems often relies on subjective user feedback. A rigorous, quantitative comparison between these paradigms remains lacking. This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages. It focuses on two key aspects: the ranking accuracy for keyword-based systems and the completeness of retrieved information for semantic chat-based systems. Our approach enables semi-automatic comparisons of semantic and keyword-based methods using interchangeable equivalence classes tailored to domain-specific contexts (e.g., companies or problems). We validate the framework through an industrial case study, demonstrating statistically significant improvements in context-aware search over keyword-based methods, supported by analyses including the Mann-Whitney U-Test. With its adaptable design, the proposed framework provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.
No comments yet — start the discussion below.