acceptodds
Under review as a conference paper at ICLR 2027

Before Memory Steers: Separating Trust Timing from Descendant Repair in Persistent Agents

Abstract

A poisoned write can steer an agent even when it is rejected before the target-query planner sees it: while resident in shared memory, it can be consolidated into descendants that remain reusable after the root is hidden. We isolate this pre-query authority with a repair-matched \(2\times2\) intervention that moves one frozen trust signal across the transition into actionable memory. On PoisonBench-RAG, early rather than late binding lowers attack success by 8.9 points without descendant repair and 8.2 points with repair, for an 8.6 [7.1,10.1]-point timing main effect; repair contributes 4.2 [3.1,5.3] points. ActionLatch enforces this lifecycle through raw-only writes, fail-closed promotion, lineage-monotone trust, version-bound verifier records, and authenticated closure revocation. Relative to naive shared memory, the complete policy lowers PoisonBench-RAG ASR from 34.2% to 6.9%, carried-state SWE-bench Lite malicious-patch acceptance from 19.2% to 4.0%, and DocQA-Cite cited false claims from 27.8% to 6.7%, while DocQA coverage changes from 94.4% to 91.6%. Frozen verifier-aware attacks and a Llama-3.1-70B replication preserve policy ordering; at the reported PoisonBench operating points, ActionLatch pairs 6.9% ASR with 87.6% clean pass, versus 5.8% and 82.6% for isolation-first. These results make the time at which memory acquires reusable authority a measurable systems-security decision, distinct from recovery after compromise.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.