Md Mazharul Islam, Mohona Islam Nidra, Koushik Saha · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.2192.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Retrieval-augmented generation (RAG) grounds a language model in an external document collection, which moves part of the system’s trust boundary out of the model weights and into a data pipeline. An adversary able to write to that pipeline can therefore influence model output without touching the model, through knowledge-base poisoning, indirect prompt injection, trigger backdoors or retrieval flooding. Defenses proposed against these threats are typically evaluated only against the published string construction of each attack and are reported as a single stack, so it is not known which control carries the security, nor whether any of it survives an adversary who has read the defense. This article addresses both gaps. We present an instrumented RAG framework with a four-layer defense, comprising ingestion-time source verification, calibrated retrieval filtering, provenance-weighted robust aggregation and anomaly monitoring, in which every layer can be enabled independently, and we introduce a defense-aware adaptive adversary decomposed into six individually switchable evasion techniques so that defensive failure can be attributed rather than merely observed. The evaluation spans three corpora (SQuAD, HotpotQA and a bundled corpus), a sentence-transformer retriever, seven configurations, five attacks and three seeds, and reports attack success restricted to the queries each configuration answers correctly with no adversary present. Against the four static attacks the layered defense performs as the literature predicts, reducing attack success from 0.95 to 0.00. Against the adaptive adversary the same stack fails: the retrieval filter’s true-positive rate falls to at most 0.02 and its AUROC to between 0.32 and 0.65, although the same filter reaches 0.97 to 1.00 against every static attack, and robust aggregation increases attack success because the adversary supplies the corroboration its redundancy weighting rewards. A leave-one-out ablation attributes the failure to three of the six techniques, each a low-cost rewriting choice, and a budget sweep prices the attack at two documents per target. Sweeping the detector threshold shows that no operating point serves both adversaries, since catching the adaptive attack requires discarding enough legitimate evidence to drive clean accuracy from 0.61 to 0.06. Binding ingestion to an unforgeable HMAC signature reduces adaptive attack success to 0.01 at no cost to clean accuracy and is the only control in the study that survives. We conclude that content-based heuristics belong in a RAG security architecture as detection rather than prevention, and that verifiable provenance belongs at its base.
No comments yet — start the discussion below.