acceptodds
Under review as a conference paper at ICLR 2027

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

Abstract

Personalized language agents keep a persistent memory so they can adapt to a user over weeks rather than turns, and that memory is also the attack surface. When an agent reads a statement that contradicts what it already believes, it has no principled way to separate a genuine change of mind from a temporary context switch, a sarcastic aside, or an instruction smuggled in through a retrieved document. Personalization methods treat every update as authentic, memory defenses reject untrusted content without modeling how real preferences evolve, and the two failure modes therefore trade against one another. We formulate the problem as a continuous-time partially observable decision process over a latent user state, and we prove that the error of any rule reading only recency and provenance is bounded below by how closely a feasible adversary can imitate the statistics of a legitimate revision, that conditioning on the interaction history is never worse and is strictly better under a stated condition, and that a bounded-cost clarifying question enlarges the achievable error region under a stated condition. CAPTURE tracks the latent state with a neural differential equation, writes hypotheses into a ledger stratified by timescale, queries the user when its belief is too uncertain to act on, and audits the causal influence of its own citations by counterfactual removal. On 480 held-out episodes from 96 users, CAPTURE reaches a 71.5% win rate against 69.3% for a baseline given identical supervision and 66.1% for the strongest heuristic, holds fixed-policy poisoning to 11.5%, and adheres to 83.5% of genuine updates. An adaptive attacker with the released weights raises poisoning to 24.7%, at which point a provenance filter is marginally more secure and considerably less adaptive, so we report both numbers rather than the favorable one. Because a system trained against our own generator could be learning that generator, we also evaluate the frozen system zero-shot on an independently built benchmark and replay the interaction histories of 40 participants, each spanning two to three weeks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.