acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Continual Learning For Vision Language Models

Abstract

Continual learning (CL) for multimodal large language models is currently benchmarked with parameter-heavy mixture-of-experts and LoRA routing methods. A large part of the reported gains on popular CL suites comes from tasks being trivially separable, so a router only needs to identify the task, instead of transferring knowledge across tasks. To address this, we build a continual multimodal multiple-choice benchmark by splitting two standard evaluation suites: MMMU (6 course categories / 30 fine categories) and MMStar (6 course categories / 18 fine categories) into disjoint task streams. Following previous work, we quantify their separability with a nearest-prototype routing probe and demonstrate that our benchmark is not trivially separable. We also introduce COMEM (Content- based Mem), a training-free, content-based k-nearest neighbor memory for a frozen Vision-Language Model (VLM). COMEM stores, per input example, the model’s last-layer question representation as a key and question-conditioned embedding of the correct answer’s content as a value. At test time, it interpolates the model’s own answer choice distribution with a distribution obtained by matching each option to a similarity-weighted retrieved answer meaning. Because the memory is append-only and avoids modifying model weights, COMEM has essentially zero catastrophic forgetting by construction. Across all four evaluation streams, COMEM consistently improves over a frozen model near-zero backward transfer, as well as matches or outperforms the mixture-of-LoRA-experts current state-of-the-art method, all without updating any model weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.