Eda Erer, Muhammet Talha Odabaşı, Ahmet Oğuz, Mustafa Circi, Mehmet Yavuz Yağcı · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202610.0008.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
As large language models (LLMs) increasingly evolve from passive text-generation systems into autonomous agents capable of interacting with external tools, databases, file systems, and operating environments, they introduce new attack surfaces that extend beyond conventional prompt manipulation. In particular, hybrid prompt injection attacks can embed malicious natural-language instructions designed to trigger application-layer vulnerabilities, including path traversal, SQL injection, and command injection. General-purpose prompt injection detectors are often insufficient for identifying such domain-specific, obfuscated, and execution-oriented payloads. To address this limitation, this study introduces PromptWAF, a modular, multi-layer defense orchestrator designed to intercept and neutralize adversarial prompts before they reach the underlying LLM or associated execution pipeline. As a core contribution, we construct and publicly release extensive domain-specific datasets covering LLM-mediated path traversal, SQL injection, and command injection attacks, thereby providing a reusable benchmark for the systematic evaluation of security mechanisms against execution-oriented prompt threats. Using these datasets, several transformer-based architectures, including DistilBERT, RoBERTa, and DeBERTa-v3, are systematically fine-tuned and evaluated. In addition, VulnMCP, a vulnerable testing environment based on the Model Context Protocol (MCP), is developed to emulate chained exploit scenarios in agentic LLM applications. Experimental results obtained on an isolated external test set show that the proposed models substantially outperform zero-shot baselines. The final PromptWAF ensemble achieves consistently high detection performance and near-perfect recall across the targeted attack categories, demonstrating its effectiveness in blocking specialized adversarial prompts while maintaining deployment-oriented modularity. To support reproducibility, facilitate independent validation, and accelerate further research on LLM-induced application-layer attacks, we openly release the curated datasets together with the fine-tuned models and the VulnMCP evaluation framework.
No comments yet — start the discussion below.