acceptodds
Under review as a conference paper at ICLR 2027

Projected Routing In Scaled-up Mappings for Continual Post-Training

Abstract

Continual post-training of pre-trained models offers a promising way to acquire new classes while retaining a shared representation and allocating adaptation capacity to individual tasks. However, this setting introduces a fundamental inference problem: the task associated with an input is unknown, yet selecting the correct task-specific adaptation is necessary for accurate classification. Existing approaches either concatenate features from all task-specific modules, which grows the representation dimensionality, or evaluate every module under a selection criterion such as distance or predictive entropy. The latter family is effective when the correct expert is selected, yielding high task-local classification accuracy. However, expert selection remains the bottleneck: batch-level criteria assume that test samples belong to the same task, while the per-sample alternative requires evaluating every stored module, making inference cost grow linearly with the number of tasks. We introduce PRISM (Projected Routing In Scaled-up Mappings) that explicitly separates task identification from classification. It learns a globally consistent router in a high-dimensional random feature space, enabling instance-wise selection of a single task-specific adaptation module without evaluating all adapters. This design provides constant-path inference while maintaining a shared routing function throughout the continual sequence. Across standard class-incremental benchmarks, PRISM outperforms existing methods, with particularly pronounced improvements under challenging distribution shifts. These results demonstrate that explicitly decoupling task identification from class prediction provides an effective and scalable mechanism in continual post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.