Beyond Outcome Rewards: Verifiable Information Gain for Reasoning Reinforcement Learning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) improves large language model reasoning, but its outcome-level rewards provide limited supervision for intermediate reasoning steps. Recent self-distillation methods introduce teacher guidance, yet densely supervising entire teacher trajectories may inject redundant signals and constrain policy exploration. We propose VIRAL, a Verifiable Information-guided Reasoning-level Advantage Learning framework that selectively supervises critical reasoning patterns. VIRAL constructs a privileged teacher conditioned on the ground-truth solution and the student’s response, identifies informative teacher patterns via ground-truth information gain and influential student patterns via uncertainty reduction, and establishes semantic correspondence between their reasoning trajectories. For the resulting critical patterns, VIRAL measures local teacher–student discrepancy and incorporates it into the RLVR advantage, while leaving the remaining trajectory governed by the original outcome reward. Experiments across three LLMs and one VLM on eleven reasoning benchmarks demonstrate consistent improvements over strong RLVR and self-distillation baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.