Harnessing Internal States to Defend Memory-Poisoned Agents
Abstract
Memory poisoning can redirect an LLM agent’s decisions, yet behavior-level evaluation offers limited visibility into how retrieved poisoned memories shape the agent’s internal decision process. We introduce an internal-state-aware harness that uses structured internal diagnoses as feedback for defense refinement. Using Natural Language Autoencoders (NLA), we translate token-level hidden states into natural-language summaries and map them into six local diagnostic states over constraints, objectives, and actions. These states reflect which component of the agent decision is compromised rather than following a universal temporal progression. NLA-guided activation steering further shows that the diagnosed internal semantics are behaviorally relevant, but stronger interventions increasingly degrade output executability, limiting direct steering as a standalone defense. We therefore aggregate recurring diagnostic failures and use them to iteratively refine an upstream natural-language guardrail in the agent harness. Across five agent settings and four model backbones, the refined policy substantially suppresses post-retrieval attacks while preserving benign utility, reducing the compromised-state ratio by 70% and achieving an average 78-percentage-point reduction in attack success rate. Overall, our results show that internal states are most useful as structured feedback for preventing poisoned decisions before they form, rather than as direct control targets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.