acceptodds
Under review as a conference paper at ICLR 2027

Gradient-Space Diversity as Credit Assignment for Post-Training Large Language Models

Abstract

Reinforcement learning has become a standard approach for post-training large language models, but improvements in task performance often come with reduced output diversity and increasingly narrow solution patterns. Most existing diversity-aware training methods quantify variation in output, embedding, or hidden-state space and optimize the resulting metric as an auxiliary reward. In contrast, this work leverages training gradients as the primary signal to measure and promote diversity. First, under a local first-order approximation, we show that GRPO propagates group-relative advantages through the rollout-gradient Gram matrix, exposing gradient redundancy among verified solutions as a hidden dimension of credit assignment. Motivated by this observation, we introduce GraDiC, a novel training method that directly operates on this gradient space. GraDiC measures their coverage using the order-2 Vendi Score and differentiates this objective with respect to each response's update weight, yielding a closed-form influence score that quantifies the response's marginal contribution to gradient-space diversity. The resulting update mechanism promotes underrepresented gradient directions and mitigates the tendency of policy updates to concentrate along dominant directions. Experiments on challenging mathematical reasoning benchmarks show that GraDiC substantially improves pass@ for compared with standard baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.