UnL: Unlocking Label Likelihoods in Multimodal Large Language Models for Class-Incremental Learning
Abstract
Multimodal large language models (MLLMs) have shown strong potential for class incremental learning (CIL), as flexible semantic generation supports the acquisition of new classes while retaining recognition of previous ones. In this paper, we reveal overlooked inference interference induced by label-space expansion, with persistent generative preference among candidate classes progressively biasing predictions toward newly introduced classes and degrading recognition of previous ones. To address this issue, we propose Unlocking Label Likelihoods (UnL), which mitigates inference interference by calibrating preference from generative evidence. Specifically, UnL derives generative evidence over the accumulated label space through Candidate Scoring and dynamically updates a calibration queue using the derived evidence, enabling Online Calibration to estimate a class-wise correction via Sinkhorn matching for subsequent predictions. Theoretical analysis establishes exact absorption of the shared additive class-wise preference component and bounded sensitivity of calibrated pairwise evidence to sample-specific deviations. Extensive experiments on multiple CIL benchmarks show consistent improvements over existing methods, with larger gains as the number of incremental tasks increases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.