PRISM: A Probe Suite for Retrieval, Interference, and Selective Memory
Abstract
We introduce PRISM, a suite of seven synthetic probes for in-context retrieval. The probes vary how key-value bindings are arranged in the context (dispersion) and how many must be held at once (scope), with sequence length and vocabulary fixed across three difficulty tiers. We train six architectures (GDN, GDP, RWKV-7, Mamba2, MLA, Transformer) from scratch with dense next-token cross-entropy at matched parameters and training budget, with per-seed reporting and length and complexity generalisation tests. We control for the most direct state-size confound by matching nominal recurrent state dimension to the hardest binding count. We find that: (i) under this small-scale, fixed-budget regime, RWKV-7 solves nearly the entire suite and GDN/GDP solve task-dependent subsets, while Transformer, MLA, and Mamba2 solve no task; (ii) controlled interventions show that retrieval-supervision density is a major factor in whether softmax attention learns to retrieve, which helps reconcile PRISM with MQAR and MAD, while Mamba2 solves retrieval only at smaller vocabulary and shorter sequences; (iii) the structured-versus-unstructured dispersion axis dissociates models that the difficulty knob alone does not; (iv) when sequence length and the difficulty knob are each scaled by the same factor (2× or 4×), complexity generalisation degrades more than length generalisation on five of seven tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.