PRISM: Probing and Repairing Internal States in Language Model Pretraining
Abstract
Internal monitors reveal changes that aggregate training statistics obscure, but deciding when to intervene requires understanding their consequences. We present PRISM, a framework that connects component-level measurements to counterfactual probes and matched interventions in language model pretraining. Across dense and mixture-of-experts models, these comparisons reveal an intervention gap: suppressing an activation tail can leave training worse, and restoring expert opportunities can coexist with higher observed loss. Conversely, a small aggregate error in optimizer state can precede a catastrophic update. To recover from this failure, we introduce protected-summary recovery, which reconstructs damaged Adam second moments from one independently maintained total per tensor and intact momentum. A constrained reconstruction combines this total with a known moment bound; an equal-information mean reconstruction isolates the role of coordinate allocation. Both rules contain update transients and achieve lower early-horizon negative log-likelihood than moment reset in all five matched execution blocks across two dense settings. PRISM turns internal observations into experimentally testable intervention choices, with compact optimizer-state recovery as a concrete application.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.