BLOO: Bayesian Leave-One-Out Policy Optimization for Group Credit Assignment
Abstract
In group-based policy optimization, e.g., GRPO, relative credit assignment relies on contrasting sibling responses. When rollouts yield identical binary outcomes in early or difficult tasks, standard empirical normalization collapses, leaving entire groups with zero gradient signal. Yet agreement among a few observed rewards does not mean the policy has no room to improve. We present BLOO, a simple RLVR framework that treats peer outcomes as finite evidence rather than absolute baselines. By computing a Jeffreys-smoothed sibling reference and scaling credit by the predictive standard deviation of the next outcome, BLOO maintains well-behaved, non-zero credit even in homogeneous groups. Beyond ensuring non-zero updates, we provide a comprehensive theoretical characterization of the retained signal: (i) we prove that the expected update follows a well-defined transformed objective; (ii) we establish that homogeneous groups account for at most of the per-prompt mean signal coefficient; and (iii) we decompose gradient variance into cross-reward and within-reward components, revealing that credit retention does not guarantee variance reduction. Ablation analyses confirm the theoretical predictions, while comprehensive experiments on Qwen3-VL show that BLOO outperforms GRPO in both final and late-stage average validation accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.