Alibek Orynbek, Kainizhamal Iklassova, Azamat Serek · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202610.0085.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Most benchmarks used to evaluate retrieval-augmented generation (RAG) systems assume a static world: they say little about how well a system handles new information, conflicting evidence, multiple languages, or questions that simply cannot be answered from the available evidence. This study introduces VERITAS, an adaptive, temporal- and conflict-aware RAG pipeline built to improve multilingual answer quality without sacrificing source grounding. It combines need-based update routing, hybrid BM25/dense retrieval, cross-lingual evidence access, multilingual reranking, temporal and source-trust ranking, conflict verification, history-aware filtering, adaptive context construction, and citation verification, all built around a single frozen Qwen3-8B generator shared across every method compared. To evaluate it, we built Adaptive Knowledge v6, a 3,000-item benchmark spanning Kazakh, Russian, and English across five knowledge conditions: stable, new, updated, conflicting, and unanswerable. On a deterministic 999-item test split, VERITAS reached 60.06% accuracy and 74.02% token F1, against 52.95% and 70.62% for the strongest baseline, a statistically significant +7.11 percentage-point gain (95% paired-bootstrap CI [4.40, 9.71]; McNemar p = 2.22 × 10⁻⁷). The improvement was largest on conflicting and unanswerable items but did not carry over uniformly to updated knowledge. VERITAS also showed a coverage–grounding trade-off, with accepted-answer coverage of 70.67%. Together, these results point to real gains under heterogeneous knowledge conditions, while also surfacing wrong abstention and updated-knowledge handling as the clearest targets for future work.
No comments yet — start the discussion below.