Measuring What Must Persist: A MiniGrid Diagnostic of State-Formation Scaffolding
Abstract
Partially observable RL agents must decide what information should persist as internal state. We treat this as a measurement problem on a controlled MiniGrid cue_recall diagnostic, using a slot module (ASF) as an instrument: keep/forget gates, cue-decode readouts, and causal memory interventions—not as a general-purpose memory architecture. Unless noted, training-run success is the last 100 episodes of a 50k-step PPO log (3-seed); frozen-policy and distillation numbers are 200-episode evals. With privileged cue-identity losses and event-gated writes, ASF reaches 84.7% vs. 29.7% feedforward PPO and 35.0% GRU; the training auxiliary cue-head decodes the cue from memory at 100%, and zero-masking memory at decision drops success by 21–64pp across privileged Claims A–D. The same template, with delay, distractors, and slot budget varied one at a time, yields 262 selected runs in the submission inventory across four memory pressures. Retention loss (−34.3pp at b=40) and cue-identity auxiliary losses (−11.3pp vs. Claim A full) are load-bearing; the compression loss is not. Removing both cue losses and privileged writes collapses ASF to 3-way chance (30.7%), alongside a retuned GRU (31.3%): the 84.7% result measures scaffolding, not slots versus recurrences. Replacing simulator cue_id with a pixel readout—no probe labels at train or test—yields 90.0% from-scratch ASF. Distilling that student's write vector, with no discrete cue_id at test and no privileged teacher, is within 0.3pp of a 200-episode eval of the student (92.8% vs. 92.5%; n=3, not an equivalence test). Writing observation tokens remains at 3-way chance. The load-bearing object is an identity-shaped write channel; PPO does not invent it (23.0% from-scratch write-head). Under our protocols the identity write channel is required for the observed gains, but its effect is architecture-dependent and high-variance: the same identity write yields 72.7%±13.4 (88/63/67) in a Transformer (w=16) vs. 37.0% in a GRU. We report this isolation rather than claiming a new memory module.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.