acceptodds
Under review as a conference paper at ICLR 2027

Multimodal RLVR With Routed Distillation

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves multimodal reasoning, yet a shared outcome-derived advantage cannot distinguish errors in image perception from errors in reasoning over the image. These functions often recur within the same response, making both the location and the source of expert feedback consequential. We introduce RLRD, short for RLVR with Routed Distillation, which allocates token-level feedback from frozen perception and reasoning experts with different post-training backgrounds. Image-perturbation sensitivity and predictive entropy serve as routing proxies, selecting positions along each on-policy response for the corresponding expert. The routed teacher-student log-probability gaps are incorporated into an existing sign-preserving advantage reweighting rule, adjusting local update magnitudes while retaining the update direction set by the verifier reward. Across six vision-language benchmarks and Qwen2.5-VL-3B/7B students, RLRD achieves the highest perception and reasoning averages among the evaluated methods, improving over the strongest non-RLRD baseline in each group by 1.39-2.02 percentage points. Controlled ablations show gains over random routing at the same teacher-token budget and over swapped or averaged expert assignments. Together, these results support selecting both feedback positions and expert sources to improve learning from interleaved multimodal responses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.