Selective Computation Declines After Short-Answer Continual Fine-Tuning: Elicitation and Target Format
Abstract
Standard continual learning (CL) metrics measure forgetting strictly within the new task stream, leaving a massive blind spot: the degradation of a pretrained model's prior, general capabilities. Here, we uncover a striking, highly selective failure mode induced by continual adaptation. Fine-tuning LLMs on a classification stream severely degrades their accuracy on out-of-distribution queries requiring computation, while leaving retrieval queries almost intact. Concurrently, adaptation triggers a drastic shift in stopping behavior, forcing the model to terminate immediately after brief answers. While this stopping shift and the computation loss appear causally linked, we rigorously decouple them: manipulating only the training target format completely reverses the stopping policy, yet the computational drop remains unchanged. Crucially, the underlying capability is not erased. Given a Chain-of-Thought (CoT) scaffold, the adapted model successfully recovers correct, step-by-step derivations for queries it previously failed to answer autonomously. We conclude that continual adaptation suppresses the model's intrinsic propensity to reason rather than destroying its mechanics. The model remains fully capable but becomes fundamentally unwilling to derive answers—a silent deficit standard CL metrics fail to capture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.