acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Recoverability: Suppression-Like vs. Overwrite-Like Capability Loss in LLM Post-Training

Abstract

Post-training improves task performance but can degrade general capabilities such as instruction following, factuality, and safety. The same accuracy drop can conceal two different states: a capability may remain accessible but difficult to elicit, or it may require substantial relearning. We call these suppression-like and overwrite-like forgetting. We introduce TRIAGE, a diagnostic and repair framework that combines recovery dynamics, hidden-state readability, and lightweight-repair operability into a continuous Forgetting Typing Score (FTS). FTS achieves AUROC under leave-one-injection-type-out evaluation and transfers to new suppression mechanisms, a second model family, and a second training objective. At matched mathematical-reasoning gains, RLVR leaves more accessible losses than SFT, with FTS-weighted suppression indices of and . A controlled study reproduces of this gap within supervised training by combining on-policy data with proximity-to-base regularization; realized model–base distance strongly predicts forgetting type across supervised and RL runs. FTS-guided repair turns this distinction into a practical advantage, reaching a common recovery target at – lower cost than uniform replay. The type of forgetting thus connects how post-training changes a model to the intervention that restores its capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.