CHIP: One Chi-Squared Ball for Both Errors of Reinforcement Learning in Large Language Models
Abstract
Reinforcement learning updates a large language model with data written by the model's previous version, and the error of one such update splits into exactly two terms: a bias, because the objective we can compute scores the new policy on the states the old policy visited, and a variance, because that objective is estimated from a finite batch. We find that the divergence between the new and old policies controls both errors. For ideal importance weights, equals their variance, which governs finite-batch estimation noise; the same budget also bounds the total-variation distance in the bias bound. We enforce this single budget at every sampled position through a projection layer that replaces the actor's proposed next-token distribution with the closest distribution inside the ball around the old policy. The solution has a normalized weighted harmonic form governed by one scalar, and implicit differentiation gives its gradients in closed form. The layer adds less than 1% to each training step and replaces clipping at the same point in standard training recipes. We call this algorithm CHIP (CHi-squared Projections). Replacing clipping with CHIP raises validation accuracy in all five estimator families we tested, by +0.7 to +10.7 points. In a six-scale Qwen3 ladder, CHIP leads in seventeen of eighteen cells and achieves an average validation gain of +6.9 points over the last sixty steps. On multi-turn agentic coding tasks, CHIP achieves 42.8% pass@1 on SWE-bench Verified at training step 150, outperforming CLIP by 5.8 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.