CRL-FM: Calibrated Routing for Lifelong Learning on Foundation Models
Abstract
Continual learning (CL) in natural language processing requires sequentially mastering new task distributions without eroding previously acquired capabilities. Parameter-efficient fine-tuning and prompt-based methods reduce — but do not eliminate — interference from shared parameter updates, leaving models vulnerable to representational drift as task streams grow. We propose CRL-FM (Calibrated Routing for Lifelong learning on Foundation Models), a modular CL framework built on a frozen autoregressive foundation model backbone (GPT-2), chosen to keep the framework lightweight and fully reproducible on modest compute. CRL-FM trains a disjoint Low-Rank Adaptation (LoRA) module per task and retains each in an immutable Snapshot Vault, so that later tasks cannot overwrite earlier ones by construction — yielding zero backward interference at the adapter level. At inference time, an attention-mask-aware prototype router computes normalized semantic centroids over frozen contextual embeddings to identify the target task without retaining raw exemplars, and a calibrated selective-prediction layer abstains on low-margin queries whose top-two prototype similarities are ambiguous, trading coverage for precision under a tunable threshold. On a five-task sequential benchmark (SST-2, AG News, IMDb, DBpedia-14, Yahoo Answers Topics) across four random seeds, CRL-FM's task-specific adapters attain 82.18% ± 0.22% average accuracy under known task identity, substantially outperforming Sequential Fine-Tuning (43.42% ± 4.62%) and matching or exceeding Experience Replay and L2P while incurring no adapter-level forgetting. We further characterize the end-to-end, task-agnostic setting: the prototype router achieves 78.88% top-1 routing accuracy, and we show that selective prediction recovers high reliability (98.73% routing accuracy at 21.07% coverage) by abstaining on ambiguous queries — quantifying, rather than obscuring, the accuracy–autonomy tradeoff inherent to task-free continual deployment. We position CRL-FM as a transparent, resource-efficient reference point for parameter-isolated continual learning, and discuss its relationship to routing-based PEFT methods such as SAPT and orthogonal-subspace approaches such as O-LoRA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.