STAC: Evaluator-Free Medical Multimodal RLVR via Stability-Aware Credit Assignment
Abstract
Although reinforcement learning with verifiable rewards (RLVR) has shown promise in medical multimodal large language models, binary outcome supervision cannot distinguish robust reasoning from fragile successes that arrive at the correct answer by chance. While process reward models offer finer-grained guidance, they introduce heavy annotation or external evaluator overhead. In this work, we reveal that the rollout policy’s own intermediate states inherently reflect trajectory stability: early prefix support for the eventual answer reliably predicts continuation stability under resampling, serving as a lightweight process signal orthogonal to final correctness. Grounded in this insight, we propose STAC (STability-Aware Credit Assignment), an evaluator-free framework that refines credit assignment using internal policy dynamics. STAC introduces Prefix Support Probing to evaluate early intermediate states against competing hypotheses, and couples it with terminal verification via Outcome-Conditioned Reward Modulation to reward stably supported correct reasoning while suppressing credit for stubborn errors. Crucially, STAC requires no external judges, step-level annotations, or rollout branching during training. Across six diverse medical visual question answering benchmarks, STAC consistently outperforms strong RLVR baselines, significantly improving both diagnostic accuracy and reasoning stability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.