MM-DataLoop: Closing the Data-Side Loop for Recursive Self-Improvement in MLLM Post-Training
Abstract
Recursive self-improvement (RSI) offers a path toward automated training loops for large models. Existing research focuses primarily on the model side, while data preparation remains a labor-intensive part of MLLM post-training. Data-side methods typically address individual operations, leaving the composition of heterogeneous interventions and the accumulation of cross-round experience less explicitly organized. We present MM-DataLoop, an end-to-end iteration loop for MLLM post-training data. It groups data operations into five families and runs an Observe Diagnose Plan Execute Verify Reflect cycle. A global rule memory carries experience into subsequent plans. We hold the base-model initialization, training hyperparameters, and evaluation protocol fixed while varying the training data and its mixture. This contract makes data interventions comparable and auditable across rounds. We run 30+ rounds on Qwen3.5-0.8B across MMBench, ChartQA, MMMU, MathVista, and RealWorldQA. The five-benchmark average improves from 57.72% zero-shot to 62.95% (5.32 pp). The study also yields reusable empirical findings about filtering and data composition. For example, removing different subsets flagged by one groundedness filter produces opposite effects. Some subsets judged low-quality also provide useful visual diversity, and removing them reduces performance in the tested recipes. We will release the code, data, and empirical findings to support further data-side research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.