acceptodds
Under review as a conference paper at ICLR 2027

Skill-: Skill-Policy Co-Evolution for Self-Improving Multimodal Reasoning

Abstract

Post-training with explicit reasoning trajectories is important for enhancing the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, obtaining high-quality reasoning trajectories typically requires costly human annotation or supervision from stronger external models. To reduce this cost, self-improvement paradigms have emerged, enabling models to generate, verify, and learn from their own reasoning experience. Despite their effectiveness, existing self-improvement methods remain largely trajectory-centric, repeatedly consuming sample-specific reasoning trajectories while leaving persistently unresolved samples with little contribution to subsequent improvement. In this work, we introduce Skill-, a skill-policy co-evolution framework that turns self-generated multimodal experience into reusable reasoning skills and continuously expands their capability coverage. In Skill-, a shared MLLM policy jointly performs skill selection, utilization, and evolution, with stage-specific credit assignment yielding learning signals according to the downstream consequences of each decision. The newly-recovered successful experience further improves the parametric policy and evolves the skill library, allowing reusable skills and the policy to improve iteratively across evolution rounds. Experiments show that Skill- achieves up to higher evolutionary efficiency and greater data efficiency than existing self-improvement methods. Across diverse benchmarks spanning mathematical, chart/document, and general multimodal reasoning, Skill- also achieves stronger performance against representative multimodal reasoning models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.