acceptodds
Under review as a conference paper at ICLR 2027

Recursive Continual Post-Training of Foundation Models

Abstract

Foundation models should be able to learn new tasks after deployment without forgetting old ones or becoming larger over time. Current approaches tend to preserve one at the expense of the other: self-distillation keeps the model size fixed, but is limited by the model’s own capabilities; expert routing protects past tasks with dedicated frozen experts, but every new task makes the deployed model larger and more expensive to run. We present Recursive Continual Post-Training (ReCap), which captures the benefits of frozen experts without deploying them. At each stage, ReCap trains an expert on the new task, then distills it back into a single base-sized checkpoint. For prior tasks, a small per-task generator reconstructs input embeddings, which the corresponding frozen expert relabels to recover its behavior in the student. These frozen experts provide stable distillation targets across later stages. On CoIN and UCIT-O, ReCap reaches and average accuracy, respectively, matching or outperforming the strongest expert-routing baselines. It also improves average performance across six held-out benchmarks by 2.1 points over the base model, with no benchmark dropping. Similar results hold across Qwen and Llama models from 1B to 7B, and across both vision-language and text-only task sequences.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.