CreditOmni: Outcome-Constrained Segmental Credit Assignment for Multimodal Reasoning
Abstract
Multimodal reasoning requires models to interpret audiovisual evidence and use it to support their final answers. Reinforcement learning can improve this ability with verifiable answer rewards, but these offer limited guidance on evidence and reasoning quality. In particular, responses with the same answer score may differ in the quality of their evidence and reasoning. In this paper, we propose CreditOmni, a framework with two components. (1) Outcome-constrained segmental credit assignment distinguishes responses with the same answer score by evaluating evidence and reasoning separately. Process scores adjust the magnitudes of answer-derived advantages while preserving their signs. (2) Two-stage group relative policy optimization follows supervised fine-tuning and proceeds from segmental learning to full-sequence refinement. Experiments on IntentBench, Daily-Omni, and WorldSense show that CreditOmni achieves the best overall performance among the evaluated open-source 7B models. It outperforms Qwen2.5-Omni-7B by 7.52, 13.37, and 3.24 percentage points, respectively. Our code is available at https://anonymous.4open.science/r/creditomni/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.