acceptodds
Under review as a conference paper at ICLR 2027

Rolling Fair Dice: Group-Level Credit Assignment for LLMs

Abstract

Large language models (LLMs) are widely used to sample from target distributions, yet audits show they fail goodness-of-fit tests on most standard distribution families. We turn this into a training problem with a verifiable reward: the total-variation (TV) distance between a rollout group's pooled empirical distribution and the target. This reward breaks GRPO-style training: broadcast to every member, it makes the centered advantage identically zero and the policy gradient vanish (Proposition 1). We introduce two remedies: JCA (Jackknife Credit Assignment) credits each sample with its leave-one-out contribution to the group score, a black-box estimator for any group-level reward, and a differentiable -diff objective reweights the empirical CDF by likelihood ratios. Under decomposable rewards, JCA reduces exactly to standard GRPO (Proposition 2). On 15 Bad-Dice families, JCA lifts a Qwen-4B model from 0.59 to 0.86 TV score and -diff to 0.93; a 9B model reaches 0.93±0.01, within two points of the strongest 2026 frontier model, without any test-time reasoning. Fine-tuning on samples from the target itself reaches the 0.96 ceiling, pricing what reward-only access costs; training generalizes zero-shot to held-out families (0.93) but stays anchored to list formats and does not transfer to natural-format tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.