VeriMem 2.0: Verification-Gated Declarative Memory for LLM Agents — Prevention at Write, Not Detection at Use
Abstract
LLM agents persist declarative memory—facts, rules, preferences—across episodes, but existing write policies admit entries on the model's own judgment (LLM-judge scoring; success filters), and both write hallucinations: under matched extractors and retrievers, ungated writes contaminate 12.8% and LLM-judged writes 17.9% of audited location claims on ScienceWorld (pooled over three seeds), while the best LLM agent answers memory-validity probes correctly only 55.2% of the time. We present VeriMem 2.0, a memory admission gate that requires each declarative write to carry an executable evidence predicate—a decidable check over a logged evidence corpus—and to abstain when no predicate is expressible. In a controlled comparison where only the gate varies, VeriMem reduces contamination to 1.5% on ScienceWorld (0.0% on ALFWorld) and sustains this under adversarial injection of up to 80 forged entries (dose 0.4), landing zero of them on both environments while ungated and success-filtered stores absorb 100%. It holds success rate under attack while ungated memory degrades on ALFWorld (0.35 vs. 0.15). Ablations attribute the entire defense to the executable predicate: removing it collapses VeriMem to no-gate and admits 32/40 forgeries, while removing evidence-log retraction changes nothing (0/40). The gate costs clean success rate within noise (0.140 vs. 0.113 ScienceWorld; 0.22 vs. 0.28 ALFWorld) and buys a 3–5× smaller, 100%-verified store; its abstention (35.2% ALFWorld, 32.1% ScienceWorld) is a tunable fallback ablated directly. All experiments use a 7–8B open-source backbone (Qwen3-8B). Code, prompts, and audit tooling will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.