Revisiting Continual Learning for MLLMs: A Simple Approach Reveals Hidden Limitations in Existing Benchmarks
Abstract
Multimodal large language models (MLLMs) have emerged as a promising foundation for AI systems, and continual learning has attracted increasing research attention to enable MLLMs to adapt to evolving environments. Despite the proliferation of related work, a critical question remains: do current benchmarks adequately capture the task relationships needed for effective knowledge sharing in MLLM continual learning? In this paper, we take an initial step toward answering this question. First, we show that a small modification to one existing MoE-LoRA method-removing cross terms from its fused LoRA readout-is sufficient to substantially outperform state-of-the-art methods. To understand this unexpected improvement, we analyze the routing behavior and find that task identity can be readily inferred from the input, allowing models to exploit task-level shortcuts and revealing a potential benchmark-hacking issue. To systematically examine this issue, we introduce an entropy-parameterized framework that recasts existing benchmarks into variants with controllable task overlap. As task overlap increases, performance can improve even as task identity becomes harder to infer, while cross-task knowledge sharing becomes increasingly beneficial. These findings suggest that existing benchmarks may overstate the benefit of task specialization while underestimating cross-task knowledge transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.