COVE: Concept Vocabulary Evolution for Streaming Continual Visual Instruction Tuning
Abstract
Continual Visual Instruction Tuning (CVIT) aims to enable multimodal large language models to continuously acquire new vision-language knowledge while retaining previously learned knowledge. Recent studies move toward more realistic streaming settings, where heterogeneous data arrive as interleaved and dynamically evolving distributions. In such streams, both the reuse of accumulated knowledge and the demand for additional adaptation capacity vary over time, making predefined expert structures difficult to match to the current adaptation demand. To address this challenge, we propose Concept Vocabulary Evolution (COVE), which organizes continual adaptation as an evolving vocabulary of reusable units discovered directly from incoming data. Specifically, COVE decomposes each incoming sample over the accumulated vocabulary, selectively refining reusable historical units when they sufficiently represent the sample, while reconstruction residuals drive the discovery of additional units when existing capacity becomes insufficient. Extensive experiments demonstrate that COVE improves continual performance and substantially reduces catastrophic forgetting, while further analyses show that the learned vocabulary dynamically evolves with the changing demands of the stream.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.