Jorge Cisneros-González, José A. Ondiviela, Javier Sánchez-Soriano · Street Art & Urban Creativity 2026 · 2026
DOI: 10.62161/sauc.v12.6399
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Public administrations are deploying large language model (LLM) assistants that process, summarise, classify and validate citizen-submitted documents. These copilots are exposed to indirect prompt injection: instructions hidden in manipulated documents that reach the model as if they were data. We develop a municipal aid-procedure copilot and evaluate its robustness across six attack objectives, five delivery vectors, five defences and four LLMs, with 3,000 attack evaluations and 1,000 legitimate evaluations. We combine deterministic detection with a human-validated LLM judge. Defences reduce attack success, but unevenly across objectives. They largely neutralise imperative attacks, such as request misrouting, while barely affecting summary falsification, revealing a provenance gap: the inability to determine whether an output value originates from an authoritative field or attacker-controlled content. Moreover, attack success and potential harm are decoupled: payment fraud and rule-exfiltration attacks have the greatest potential for harm despite intermediate success rates. We frame the risk within a human-in-the-loop model, where automation bias may amplify harm.
No comments yet — start the discussion below.