A Matched-Error Sensitivity Audit of Group-Relative Policy Updates
Abstract
Marginal verifier error rates leave unspecified how errors perturb a group-relative policy update. We measure this sensitivity on fixed responses using mean-token log-probability gradients in specified LoRA coordinates. An exact decomposition compares the observed distortion with two references: one preserves error counts within each prompt, and the other preserves only the global false-positive and false-negative totals. This separates response-level error association from allocation across prompts. On a frozen 12,288-generation LiveCodeBench cohort, three-test observed-to-global-expectation ratios are 0.605 [0.504, 0.723] for Llama and 0.342 [0.225, 0.510] for Qwen Instruct, using validated labels and full-response FP32 Grams. Within-prompt association is negative for Llama and uncertain for Qwen; one-test ratio intervals include one for both policies. A retrospective APPS audit shows related allocation effects, with sparse, cap-limited Thinking evidence. Diagonal Grams preserve much of the distortion ranking, and tested full-matrix features do not establish a stable selection advantage. A restricted eight-action training experiment favors prompt-correlated noise, but free-form follow-ups and a frozen 48-run GSM8K study do not confirm that benefit. The audit identifies what a matched-error contrast measures while separating local distortion, audit-selection utility, and downstream training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.