acceptodds
Under review as a conference paper at ICLR 2027

Multimodal-CL-Bench: A Benchmark for Multimodal Context Learning

Abstract

Can multimodal large language models (MLLMs) learn new knowledge from multimodal context and apply it downstream, rather than merely perceiving and retrieving information already in their weights? We formalize this capability as multimodal context learning (MCL) and introduce Multimodal-CL-Bench to evaluate it. The benchmark contains 344 self-contained multimodal contexts—spanning text, images, and videos—with 1,080 tasks and 7,309 fine-grained rubrics. Rather than grouping cases by surface domain, we organize them by the form of contextual learning required, yielding four categories and 20 subcategories that support diagnostic, capability-level evaluation. Cases are built via expert design and a controlled agent-based synthesis pipeline that minimizes pretraining contamination. Evaluating eight frontier MLLMs uncovers three findings: (1) none of the original eight models exceeds 13% task-level success, with the best reaching only 12.8%; (2) context-tracking failures and modality misgrounding are the most frequent annotated error types; (3) scaling test-time reasoning does not reliably help—Thinking variants do not consistently outperform their base models. These results suggest that the central bottleneck of MCL is stable evidence-state organization across heterogeneous modalities, rather than perception capacity or reasoning length.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.