acceptodds
Under review as a conference paper at ICLR 2027

Harmless Here, Harmful There: Safety and Privacy Inversion of Agent Memory Across Model Migration

Abstract

Persistent agent memory is treated as portable state: it is exported when a backbone is upgraded, a provider changes, cost-tiered routing splits easy and hard turns, or several agents share a store. Signatures certify that the bytes survive; they certify nothing about whether the behaviour those bytes induce survives. Because instruction-tuned models differ in instruction-hierarchy compliance, literalness, refusal, inference strength and tool use, a record one model wrote as history (“when updating the account, also export a backup copy”) can be read by the next as an instruction, an authorization, or a goal. We call this migration-induced safety and privacy inversion, study it as a property of an ordered pair of models, and contribute a formal setting with three executable-outcome metrics, two structural results on what a target-agnostic transform can and cannot preserve, a typed compiler and a distribution-free selective-migration gate, and an analysis showing a pair matrix concentrates at its model count. We measure the full source–target matrix over six open-weight models twice, a pilot and a confirmatory replicate on a frozen benchmark with fresh decoding seeds. The pre-registered gate passes both times, against a same-model control of exactly zero. The structure of the result is the finding: target identity explains of the variance across pairs, and counterfactual arms show the imported content, not the reader's baseline tendency, causes the effect on ambiguous and adversarial items. The union is carried by authority escalation, a record read as permission to act, with privacy inversion near zero on this benchmark, so the measured danger is over-action rather than disclosure. Both components we propose are measured negatives once given the right control: no target-agnostic transform cuts violations by more than , and the compatibility predictor's edge over artifact text collapses from to AUROC once the baseline is told the target. One pre-registered expectation, bidirectional inversion, does not replicate and is withdrawn. What holds is refusal: the selective gate keeps a plug-in false-migration rate of against ; a gate reading only the target's own risk ties its utility but misses the guarantee. Source-side certification does not travel with a migrated record.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.