acceptodds
Under review as a conference paper at ICLR 2027

CreditOmni: Outcome-Constrained Segmental Credit Assignment for Multimodal Reasoning

Abstract

Multimodal reasoning requires models to interpret audiovisual evidence and use it to support their final answers. Reinforcement learning can improve this ability with verifiable answer rewards, but these offer limited guidance on evidence and reasoning quality. In particular, responses with the same answer score may differ in the quality of their evidence and reasoning. In this paper, we propose CreditOmni, a framework with two components. (1) Outcome-constrained segmental credit assignment distinguishes responses with the same answer score by evaluating evidence and reasoning separately. Process scores adjust the magnitudes of answer-derived advantages while preserving their signs. (2) Two-stage group relative policy optimization follows supervised fine-tuning and proceeds from segmental learning to full-sequence refinement. Experiments on IntentBench, Daily-Omni, and WorldSense show that CreditOmni achieves the best overall performance among the evaluated open-source 7B models. It outperforms Qwen2.5-Omni-7B by 7.52, 13.37, and 3.24 percentage points, respectively. Our code is available at https://anonymous.4open.science/r/creditomni/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.