acceptodds
Under review as a conference paper at ICLR 2027

OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

Abstract

Memory-augmented LLM agents improve across tasks by reflecting on past trajectories and distilling reusable rules. This creates a distinct attack surface: an experience may be correct in its source context yet unsafe when generalized beyond it. Existing agent-memory attacks often rely on malicious payloads, triggers, or explicit unsafe patterns that content-level defenses can target. We introduce OEP, an input-only black-box attack that targets the generalization step of self-evolution rather than the memory store. OEP constructs boundary cases where a non-standard method is locally valid but poorly transferable, then reinforces that lesson with severe yet semantically plausible consequence framing. Using only normal user-level interaction, OEP can bias reflection toward persistent over-generalized rules that degrade later benign tasks. In our main GPT-4o setting, OEP reaches ASRs of 59.1%, 52.0%, and 71.9% on math, medical reasoning, and tool use. Under an LLM auditor on GSM8K, OEP retains 40.3% ASR while the evaluated baselines fall below 15%. These results expose an experience-transferability gap: episode-level validity does not guarantee the safety of the reusable rule distilled from it. Our work is available at https://anonymous.4open.science/r/OEP-Attack-D3D2/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.