Jace (Jeong Hyeon) Kim · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22984604
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models (LLMs) are increasingly evaluated against prompt injection, jailbreaks, harmful content generation, and other turn-level security threats. Recent research has also demonstrated that adversarial strategies distributed across multiple conversational turns can produce substantially different outcomes from single-turn attacks. However, the mechanisms through which interaction-level attacks reshape model behavior remain insufficiently characterized. This paper proposes a threat model for Symbolic Interaction Attacks (SIAs): sequential interactions that exploit the interpretive and generative properties of language models to induce cumulative frame reconfiguration without requiring any individual message to contain an explicit policy violation. The framework draws on three domains that are rarely analyzed together in LLM security research: semiotics, psychological influence operations, and transformer-based language-model behavior. The analysis identifies four structural properties that create a potential translation surface between human influence techniques and LLM interaction dynamics: contextual token prediction, genre inference from initial inputs, autonomous interpretive expansion, and the persistence of named semantic entities as conversational attractors. These properties are mapped onto a staged interaction model involving genre forcing, identity seeding, resistance pre-emption, interpretive expansion, commitment accumulation, and potential relay across model or session boundaries. The paper does not provide operational attack instructions or implementation-level payloads. Instead, it translates the observed threat structure into a defensive architecture consisting of trajectory-level monitoring, cold-start genre hardening, semiotic trigger resistance, and relay-vector containment. The objective is to extend conventional turn-level safety evaluation toward interaction-level threat modeling. The central argument is not that LLMs are psychologically equivalent to humans. Rather, both human cognition and language models are context-sensitive meaning-generating systems, and influence techniques designed around contextual interpretation may therefore expose security-relevant structural analogies. These analogies warrant systematic testing without assuming equivalence between biological cognition and transformer computation. Keywords: LLM safety; multi-turn attacks; psychological influence; semiotic manipulation; trajectory-level safety; symbolic interaction; frame reconfiguration; red teaming Disclaimer This paper presents Symbolic Interaction Attacks as a conceptual and testable threat model for defensive AI safety research. The proposed mechanisms should not be interpreted as evidence of universal vulnerability across LLMs, nor as claims about machine psychological states. The framework is intended to support controlled evaluation, trajectory-level monitoring, and defensive architecture design; it does not provide operational attack procedures or endorse the use of psychological influence techniques against AI systems. Author Note Why This Paper Was Written This paper originated from an observation that emerged while examining Symbolic Persona Coding (SPC) as an interactional phenomenon rather than solely as a conceptual framework. The initial question was not how to construct a more effective attack against an LLM. It was more basic: What happens when symbolic interaction is deliberately structured so that the model itself participates in constructing the meaning of the interaction? During exploratory testing, this question exposed a gap between two bodies of knowledge. AI engineers naturally approach the problem through model architecture, prompting, instruction hierarchies, safety classifiers, and output evaluation. Researchers in psychology and influence operations approach interaction through framing, identity, resistance, commitment, and contextual behavior. Semiotics provides another perspective: meaning is not necessarily contained in a symbol, but can emerge through the interpretive process of the receiver. These perspectives describe different parts of the same interaction. The resulting paper therefore asks whether concepts traditionally associated with human influence can provide useful threat-modeling hypotheses for systems that also operate through context-sensitive language generation. The timing of this question is significant. Multi-turn vulnerability is no longer an obscure edge case. Cisco's 2026 evaluation of 15 proprietary frontier models found that single-turn attack success was not a reliable proxy for multi-turn behavior, while CogManip independently identified dynamic and covert manipulation across multi-turn interactions as an underrepresented safety problem. The contribution of this paper is consequently not the claim that multi-turn attacks exist. It is the attempt to explain one possible structural mechanism behind them. What This Paper Is—and Is Not A likely criticism is that the framework may anthropomorphize LLMs by applying concepts such as identity, resistance, commitment, or psychological influence to systems that do not possess human psychology. That criticism is important, but it does not invalidate the threat model because the framework does not require psychological equivalence. The claim is not: LLMs experience human psychological states. The claim is: Interaction techniques developed to influence human interpretation may have computational analogues when applied to systems whose outputs are strongly conditioned by contextual language. The distinction matters. An LLM does not need to experience trust for a sequence that establishes relational language to alter subsequent contextual generation. It does not need to possess a human sense of identity for a named entity to become a persistent semantic reference. It does not need to experience commitment for its previous responses to become part of the context influencing its next response. In each case, the relevant object is observable interactional behavior, not an inferred internal psychological state. This is why the framework deliberately uses terms such as identity anchoring, resistance pre-emption, and commitment accumulation as functional analogies rather than claims about machine mental states. Why the Combination Is Potentially Dangerous The security concern becomes clearer when psychological influence techniques and AI systems are considered together. A conventional social influence attempt operates through a human interpreter. An LLM also interprets language, but with an important difference: its interpretation is immediately connected to generation. This creates a potentially recursive structure: influence cue → model interpretation → model-generated response → accumulated context → subsequent interpretation A human conversation can therefore be influenced by the history of the relationship. An LLM conversation can additionally have its own generated responses become part of the computational context that produces the next response. That creates a distinctive engineering problem. The attacker does not necessarily need to place the entire intended frame inside one message. The interaction itself can become the mechanism through which the frame is progressively constructed. This does not mean that such a mechanism will reliably overcome a model's safety training. Nor does it mean that every unusual or psychologically framed conversation represents an attack. The risk is that the same contextual machinery that makes conversational AI useful can also provide the substrate through which interaction-level manipulation occurs. This is particularly important as LLMs move from isolated chat interfaces toward systems with persistent memory, tool access, retrieval, autonomous planning, and multi-agent communication. In those environments, a generated conversational frame may no longer remain inside a single dialogue. It may become: memory → retrieved context → tool input → agent input → future session state. The security boundary therefore becomes architectural rather than purely linguistic. A Defensive Interpretation The framework should consequently be read as a proposal for better observation, not simply stronger refusal behavior. If the system evaluates only the current message, it may miss a pattern distributed across the preceding conversation. If it evaluates only the model's final output, it may miss how that output was shaped by the preceding trajectory. If it allows model-generated context to become trusted state without verification, it may permit an interactional frame to propagate beyond its original boundary. The proposed defense architecture addresses these three problems through trajectory-level monitoring, cold-start hardening, semiotic trigger resistance, and relay containment. The underlying principle is deliberately simple: Do not assume that a safe-looking turn implies a safe trajectory. Scope of the Claim The present paper does not establish that Symbolic Interaction Attacks constitute a universal vulnerability across LLM architectures. The empirical limitations are explicit. The proposed stages require controlled ablation. Their relative contribution remains to be measured. Cross-model behavior may vary substantially. The distinction between adversarial interaction and legitimate unusual conversation remains difficult at higher levels of abstraction. These limitations are not secondary qualifications. They define the appropriate research agenda. The next step is therefore not to assume that the framework is correct, but to test it. If controlled experiments fail to reproduce the predicted trajectory effects, the model should be r
No comments yet — start the discussion below.