acceptodds
Under review as a conference paper at ICLR 2027

How Foundation Skills Boost Continual Vision-Language Learning: A Skill-Evolution Perspective

Abstract

Continual learning can empower large vision-language models (VLMs) to acquire new knowledge in non-stationary environments without replaying historical downstream-task data, yet continual post-training often causes catastrophic forgetting. We argue that such forgetting also stems from the degradation of reusable foundation capabilities inherited from pretraining. This paper therefore revisits a fundamental yet underexplored problem: how can foundation skills in VLMs boost continual vision-language learning performance? We propose a Skill-Centric Adaptation model for Lifelong Evolution, termed SCALE, which reinterprets pretrained VLMs as a repertoire of reusable foundation skills. As continual downstream tasks come, downstream task adaptation follows supervised fine-tuning (SFT), while SCALE adopts skill-specific drift-aware gates to identify drifting foundation skills and task-prioritized co-evolution to maintain them. These selected skills can evolve with the model through verifier-guided behavioral supervision, without constraining current-task adaptation. This co-evolution strategy allows foundation skills to better support new-task learning while mitigating forgetting associated with foundation-skill drift during task-biased post-training. A fixed maintenance budget is dynamically allocated across foundation skills according to capability drift, independent of downstream task history. We evaluate SCALE on VQA-v3 under the SS and PS settings of UCo-VQA, with analyses of continual-learning performance, generalization, sensitivity, ablations and foundation-skill evolution. Our SCALE improves final average accuracy by 8.7% and reduces final forgetting by 9.1% points over matched continual SFT.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.