Deterministic Addresses Constrain Reads, Not Ordinary Writes: Deletion Boundaries in Conditional-Memory Language Models
Abstract
Conditional-memory language models retrieve from deterministic n-gram-addressed tables, making table state directly inspectable and removable. Yet an address identifies where a model reads, not where gradient-based adaptation writes. Unlike causal-tracing locations, these addresses are computed by the architecture rather than estimated. We test whether they form a deletion boundary by grafting exact and hashed n-gram memory into pretrained Pythia (410M, 1.4B) and Qwen2.5 (0.5B, 1.5B), fine-tuning rare trigger–continuation mappings at preregistered exposure counts, and intervening on rows, grafts, and backbones. Across both families, behavior reaches near-ceiling success when the graft is frozen. In Pythia, symmetric swaps localize the learned mapping to the backbone: a clean graft preserves it, a poisoned graft on a clean backbone yields zero success in all ten swaps, and target-row deletion is inert. The same deletion produces model-level mean target-specific attack-success reductions of 98.3 to 100 percentage points for writes deliberately confined to addressed rows. In a separate preregistered durability assay, deleted mappings remain at zero observed exact-match success through 512 later optimization steps, retained Adam state, and one constructed exact 16-row collision. A paired four-rate sweep shows optimizer policy can move causal dependence into the table at a sharp, scale-dependent threshold. At the largest rate, whole-table state exceeds the registered 0.15 attack-success sufficiency contrast in 13/16 1.4B and 11/16 410M runs, both observed majorities; only 1.4B passes the registered majority-confidence rule. Clean perplexity rises 8.0% at 410M and 7.1% at 1.4B, while final-row sufficiency exceeds 0.15 in only 7/16 runs at each size. Our scope is retrofitted 0.4B–1.5B models and rare-string, single-token probes. Deterministic addressing constrains reads but does not by itself create an auditable deletion boundary; that requires an enforced write path and post-adaptation location verification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.