Trajectory-Probing Continual Post-Training: Identifying and Closing the Readout Gap
Abstract
Continual learning with foundation models is dominated by methods that adapt pre-trained representations per task via prompts, adapters, or low-rank updates, under the assumption that feature adaptation is the primary driver of performance. We revisit this assumption and find that the picture is inverted: as tasks accumulate, a linear joint probe fit on the continually adapted backbone climbs toward joint-training while the incrementally trained classifier accuracy falls. We call this discrepancy the readout gap - the share of linearly decodable information that the incremental classifier fails to recover. The gap survives a closed-form readout that is invariant to task order, so it is not reducible to classifier bias. Representations are therefore not the bottleneck; the decision layer is. This does not imply that adaptation is unnecessary, but that task-specific adaptation is incompatible with an accumulated readout: it perturbs the feature geometry the classifier depends on, whereas shared, task-agnostic calibration improves features without destabilizing it. Guided by this, we introduce Scale-and-Shift Knowledge Injection (SASKIA), which calibrates representations at two granularities that are both shared across tasks: a global scale-and-shift module aligning the backbone to the target domain, and an instance-level hypernetwork producing per-sample corrections, refined online by EMA-based model merging rather than task-specific growth. New classes are absorbed by closed-form ridge regression in this progressively refined space, requiring no rehearsal or explicit task boundaries. SASKIA attains the smallest readout gap among all compared methods and matches or outperforms them across benchmarks, using under 1% of the backbone's parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.