Skill-: Skill-Policy Co-Evolution for Self-Improving Multimodal Reasoning
Abstract
Post-training with explicit reasoning trajectories is important for enhancing the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, obtaining high-quality reasoning trajectories typically requires costly human annotation or supervision from stronger external models. To reduce this cost, self-improvement paradigms have emerged, enabling models to generate, verify, and learn from their own reasoning experience. Despite their effectiveness, existing self-improvement methods remain largely trajectory-centric, repeatedly consuming sample-specific reasoning trajectories while leaving persistently unresolved samples with little contribution to subsequent improvement. In this work, we introduce Skill-, a skill-policy co-evolution framework that turns self-generated multimodal experience into reusable reasoning skills and continuously expands their capability coverage. In Skill-, a shared MLLM policy jointly performs skill selection, utilization, and evolution, with stage-specific credit assignment yielding learning signals according to the downstream consequences of each decision. The newly-recovered successful experience further improves the parametric policy and evolves the skill library, allowing reusable skills and the policy to improve iteratively across evolution rounds. Experiments show that Skill- achieves up to higher evolutionary efficiency and greater data efficiency than existing self-improvement methods. Across diverse benchmarks spanning mathematical, chart/document, and general multimodal reasoning, Skill- also achieves stronger performance against representative multimodal reasoning models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.