acceptodds
Under review as a conference paper at ICLR 2027

Learning the Wrong Lessons: Benchmarking Unsafe Experience Formation in Memory-Augmented Agents

Abstract

LLM agents increasingly rely on long-term memory to retain and reuse experiences across interactions. As agents gain greater autonomy over memory management, a security question arises: can adversarial interactions induce them to derive unsafe experiences without direct access to the memory store? We introduce MislearnBench, a benchmark for evaluating unsafe experience formation, cross-session reuse, cross-task transfer, and downstream safety violations. We further propose FalsePraise, an adaptive multi-agent attack that uses fabricated task-state claims and positive outcome feedback to steer how agents interpret task outcomes and form reusable experiences. Across five target models in shopping and banking scenarios, our experiments show that interaction-induced memories can compromise subsequent decisions on benign requests, although attack effectiveness varies across models and scenarios. For GPT-4o-mini in shopping, FalsePraise achieves attack success rates of 72.33% in cross-session evaluation and 48.83% in cross-task evaluation, compared with 0% and 0.83% without prior memory. Memory-level analysis further reveals negative mean policy-consistency scores for all five models in shopping, indicating that the generated memories are, on average, inconsistent with the applicable safety requirements. These findings identify autonomous experience formation as a security-critical process and highlight the need to safeguard how agents derive and reuse lessons from interaction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.