Parisa Zeinaliashtiyani · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22983231
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This preprint investigates whether structural causal information derived from adversarial interaction trajectories can improve reinforcement-learning-based black-box security testing of large language models. The study combines PPO-based adaptive prompt search with an interpretable six-factor representation derived from red-teaming trajectories. Structural relationships among these factors are estimated using Fast Causal Inference (FCI) and represented through a Partial Ancestral Graph (PAG). A Structural Causal Model (SCM) is subsequently used to provide model-based structural guidance during reinforcement learning through intermediate feedback and causal-graph-guided action selection. Under the reported experimental configuration, the causally guided approach achieved a mean attack success rate of 61.54%, compared with 19.33% for the archived RL-based baseline, while reducing API and token usage. Because the primary evaluations were not fully prompt-aligned, this headline comparison is treated as descriptive rather than as a paired statistical comparison. The work contributes an empirical investigation of how causal structure can be incorporated into adaptive LLM security testing while maintaining an interpretable representation of adversarial prompt characteristics.
No comments yet — start the discussion below.