acceptodds
Under review as a conference paper at ICLR 2027

COVE: Concept Vocabulary Evolution for Streaming Continual Visual Instruction Tuning

Abstract

Continual Visual Instruction Tuning (CVIT) aims to enable multimodal large language models to continuously acquire new vision-language knowledge while retaining previously learned knowledge. Recent studies move toward more realistic streaming settings, where heterogeneous data arrive as interleaved and dynamically evolving distributions. In such streams, both the reuse of accumulated knowledge and the demand for additional adaptation capacity vary over time, making predefined expert structures difficult to match to the current adaptation demand. To address this challenge, we propose Concept Vocabulary Evolution (COVE), which organizes continual adaptation as an evolving vocabulary of reusable units discovered directly from incoming data. Specifically, COVE decomposes each incoming sample over the accumulated vocabulary, selectively refining reusable historical units when they sufficiently represent the sample, while reconstruction residuals drive the discovery of additional units when existing capacity becomes insufficient. Extensive experiments demonstrate that COVE improves continual performance and substantially reduces catastrophic forgetting, while further analyses show that the learned vocabulary dynamically evolves with the changing demands of the stream.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.