acceptodds
Under review as a conference paper at ICLR 2027

When Experts Are Not Experts: Multimodal Continual Instruction Tuning with Certified Capabilities

Abstract

Multimodal Continual Instruction Tuning (MCIT) aims to let Multimodal Large Language Models (MLLMs) acquire new skills after deployment without forgetting previous ones. Modern MCIT methods build on Mixture-of-Experts (MoE), adding lightweight experts and guarding them against interference. We revisit this design and find that the defenses tax plasticity without securing the expertise they protect. Learned routers barely distinguish tasks even when trained jointly, so the defenses guard a separation that never forms. Meanwhile, the difficulty of the stream concentrates in a few capabilities that a shared model fails to hold. Based on these observations, we introduce Union-of-Capabilities (UoC), a paradigm in which a generalist, retrained on the stored data at each refresh and frozen in between, grows only by certified capability modules. Building on UoC, we propose C2R (Certify to Route), which realizes the paradigm in two stages. (1) At training time, CERTIFY admits an expert only on verified gain over the generalist, and (2) at inference time, ROUTE dispatches a query only on certified evidence through a training-free certificate cascade, without a task label. Every installed module is trained once and then frozen, so what it has learned is never overwritten. On CoIN and UCIT, \ours sets a new state of the art, outperforming the strongest previous method with zero forgetting, fewer trainable parameters, up to 1.7 faster training at a single refresh, and 1.8 faster inference, and it keeps this lead when updated after every task, with forgetting within 0.2 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.