What Is a Conversation Made Of? A Reusable Semantic–Relational State Across Safety Tasks
Abstract
Multi-turn safety prediction tasks in an agentic system all rely on the same underlying state that accumulates over a conversation: what has been introduced and remains salient, what the participants are trying to accomplish, what restricts them, and how these relate. Systems for these tasks are typically trained and evaluated one task at a time, even though multi-task and frozen-backbone architectures are common more broadly. We show that a single lightweight conversational state can be induced once, frozen, and reused across independent safety tasks and across the distribution shifts each task encounters, without retraining the representation. A shared mention encoder (M1) and a graph inducer (M2) together extract open-vocabulary entity, action, goal, and constraint nodes with persistent identity and typed relations; this M1+M2 inducer is trained once (419M parameters), frozen, and read by three small (sub-million-parameter) readers at the granularity of a whole trajectory, a single proposed action, and a single turn boundary. Our central evidence is controlled transfer under distribution shift rather than a single benchmark score. For multi-turn jailbreak detection we show that the standard generic-benign split admits a large topical shortcut (a first-turn topic probe alone reaches 0.87 AUROC), and then hold the reader to topic-matched benign twins, naturally occurring sensitive conversations, over-refusal cases, and three unseen attack families; on the twins the frozen structured state reaches 0.84 AUROC against 0.71 for a capacity-matched dense reader over the same jointly trained encoder. For agent action authorization, a single operating point selected only on InjecAgent transfers to a live same-harness AgentDojo evaluation at 6.1% targeted attack success and 74.2% benign utility (and 1.9% / 70.9% at a high-security threshold), and remains effective across suites, attacks, and a change of agent model. The representation is frozen throughout: across held-out hard negatives, attack families, live agent environments, and agent-model changes, the same state is reused without representation retraining. Controls rule out shortcut and encoder-only explanations, and matched occlusions and retrains show each task draws on the semantic and relational channels of the state to a different degree.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.