acceptodds
Under review as a conference paper at ICLR 2027

Closing the Curation Loop: Semantic Compression and Learning-Guided Expansion for MLLMs Mid-Training

Abstract

Efficient mid-training of multimodal large language models (MLLMs) requires both reducing repetitive supervision and adapting training data to the model's changing learning needs. We propose a closed-loop data curation framework for image–text mid-training that combines coverage-preserving semantic compression with learning-guided expansion. First, we jointly assess image–text pair specificity and local semantic diversity in an MLLM-derived representation space, and use cluster-level budget allocation to construct an initial image–text training set with broad semantic coverage and reduced redundancy. We then evaluate the model's prediction loss across semantic regions using a fixed probe set, with reference-relative loss gaps serving as an empirical signal of learning potential to guide subsequent data acquisition. By alternating between model diagnosis, data acquisition, and continued mid-training, the framework adapts the training data mixture to the model's learning state while preserving semantic coverage. Under the same sample budget, our framework achieves better overall performance than existing data-selection methods across multiple widely used multimodal image–text understanding benchmarks. Ablation studies further support the complementary roles of semantic compression and learning-guided expansion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.