Diagnosing Recoverability: Suppression-Like vs. Overwrite-Like Capability Loss in LLM Post-Training
Abstract
Post-training improves task performance but can degrade general capabilities such as instruction following, factuality, and safety. The same accuracy drop can conceal two different states: a capability may remain accessible but difficult to elicit, or it may require substantial relearning. We call these suppression-like and overwrite-like forgetting. We introduce TRIAGE, a diagnostic and repair framework that combines recovery dynamics, hidden-state readability, and lightweight-repair operability into a continuous Forgetting Typing Score (FTS). FTS achieves AUROC under leave-one-injection-type-out evaluation and transfers to new suppression mechanisms, a second model family, and a second training objective. At matched mathematical-reasoning gains, RLVR leaves more accessible losses than SFT, with FTS-weighted suppression indices of and . A controlled study reproduces of this gap within supervised training by combining on-policy data with proximity-to-base regularization; realized model–base distance strongly predicts forgetting type across supervised and RL runs. FTS-guided repair turns this distinction into a practical advantage, reaching a common recovery target at – lower cost than uniform replay. The type of forgetting thus connects how post-training changes a model to the intervention that restores its capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.