Definitely Maybe: Can Multimodal In-Context Learning Induce Visual Concepts?
Abstract
Multimodal large language models (MLLMs) can adapt their predictions to multimodal demonstrations in context, but whether they can induce visual concepts from a few demonstrations during multimodal in-context learning (MICL) remains unclear. Existing evaluations often confound concept induction with recognition of familiar semantic associations. It is therefore difficult to determine whether demonstrations cause MLLMs to acquire new information or only change how existing information is accessed. We study this question in a controlled setting and disentangle these two possibilities. Specifically, we decompose MICL performance into signal power, measuring how separable the target concept is inside the low-dimensional subspace the frozen language head can see, and readability, measuring how effectively the head aligns with that separation. Across our analysis, we find that MLLMs can induce visual concepts from context. However, demonstrations primarily improve the readability of pre-existing visual signals by steering them toward the frozen readout, rather than substantially changing what the query representation encodes. Additionally, we find that this steering is not selective: demonstrations can also align irrelevant or competing concepts with the readout, creating interference that limits concept induction. Our findings suggest that current MLLMs primarily learn to read rather than write during MICL: demonstrations largely teach the model how to access information already encoded in its representations, rather than providing new task-relevant information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.