acceptodds
Under review as a conference paper at ICLR 2027

BLOO: Bayesian Leave-One-Out Policy Optimization for Group Credit Assignment

Abstract

In group-based policy optimization, e.g., GRPO, relative credit assignment relies on contrasting sibling responses. When rollouts yield identical binary outcomes in early or difficult tasks, standard empirical normalization collapses, leaving entire groups with zero gradient signal. Yet agreement among a few observed rewards does not mean the policy has no room to improve. We present BLOO, a simple RLVR framework that treats peer outcomes as finite evidence rather than absolute baselines. By computing a Jeffreys-smoothed sibling reference and scaling credit by the predictive standard deviation of the next outcome, BLOO maintains well-behaved, non-zero credit even in homogeneous groups. Beyond ensuring non-zero updates, we provide a comprehensive theoretical characterization of the retained signal: (i) we prove that the expected update follows a well-defined transformed objective; (ii) we establish that homogeneous groups account for at most of the per-prompt mean signal coefficient; and (iii) we decompose gradient variance into cross-reward and within-reward components, revealing that credit retention does not guarantee variance reduction. Ablation analyses confirm the theoretical predictions, while comprehensive experiments on Qwen3-VL show that BLOO outperforms GRPO in both final and late-stage average validation accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.