Auditing Permanence: Exact Write Accounting and a Matched-Volume Control for Optimizer-State Gates
Abstract
Gates that decide which updates reach a model's base weights are judged by whether the gated model forgets less. That comparison mixes which updates the gate selects with how many it lets through. We separate the two. A masked AdamW makes every routed write exact and logged. In runs that train only the shared parameters, all 440,401,920 routed coordinates of a 0.6B student end bit-identical to the pretrained weights. The same logging lets us build a matched-volume random control. We apply both to a controller that reads its own Adam-style moments and, per block, commits an update to the base weights, defers the example for replay into low-rank adapters, or withholds it. On a seven-phase continual-distillation stream, with a protocol fixed before the final runs, three findings hold at the level of the three stream seeds. First, selection adds nothing measurable: committing random blocks at the controller's rate, with its own deferral and replay, forgets the same (seed-level 90% interval [-0.014, +0.016]). Second, forgetting follows write volume only coarsely. Writing one to two orders of magnitude more, by committing everything or removing the agreement gate, raises forgetting in every seed. Within the gated range, neither authorized nor realized base-weight change orders the arms. Across 49 configurations the best single predictor counts adapter writes too (Spearman 0.49, against 0.21 for base writes alone). Third, the controller's permanent writes neither help nor measurably hurt retention compared with training only the shared parameters. They improve only a pre-declared secondary metric, in-phase learning (all three seeds). We also show that exact isolation blinds any gate that needs positive evidence from the optimizer's own zero-initialized state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.