What to Trust, What to Learn: A Grounded Learning Curriculum for Self-Evolving Large Multimodal Models
Abstract
In the self-evolving paradigm of large multimodal models, a Proposer and a Solver derived from a shared backbone cooperate in a closed training loop: the Proposer generates questions from visual input, and the Solver's most frequent answer is treated as a self-defined “golden label”. This avoids the need for human labels but raises a fundamental question: how is self-evolution steered, and can its self-rewarding signal be trusted? We identify two failure modes: Proposer bias, where question rewards favor invalid or unhelpful tasks, and Solver bias, where agreement-based supervision reinforces shared errors as self-generated ground truth. Across seven checkpoints on POPE and HallusionBench, unanimous errors account for 14.5–76.7% of wrong majority answers, while entropy-targeted question rewards can further increase invalid question generation. We propose a grounded learning curriculum that separates what to trust from what to learn: a frozen detector grounds candidate questions in visual evidence and enables the construction of supported answer sets in place of Solver consensus; the Solver learns from the full supported set, while the Proposer earns credit only when matched trial updates yield greater Solver progress on separate probe images. No task-specific question–answer labels, teacher model, or answer judge are used, preserving the self-evolving setting while grounding its training signals directly in visual evidence. On eleven standard benchmarks, our curriculum raises Qwen3.5-9B's average accuracy from 75.58% to 78.61%, with the largest gains on CV-Bench () and MMBench (). The same curriculum improves Qwen2.5-VL-7B and Qwen3-VL-8B, with average gains of and points, respectively, showing consistent benefits across model backbones. Evidence-grounded targets outperform majority-vote targets by 2.43 points, and a learned Proposer provides an additional 1.30-point gain over a frozen one.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.