When Self-Reinforcement Ends Self-Improvement
Abstract
Self-improvement systems use a judge to select revisions and reinforce those that pass. The quality of the judge is typically assessed through evaluation accuracy, but this metric does not capture whether errors recur on the same inputs. We show that whether the judge's errors are deterministic determines whether reinforcement ends improvement. In a bounded utility model with a deterministic binary judge and expected losses exceeding expected gains, we prove that every fixed acceptance rule that initially improves utility eventually produces negative drift. The preference for positive outputs that enables improvement also reinforces the judge's errors. For adaptive rules and judges with finitely many outputs, we derive a tradeoff between the number of revisions and cumulative negative drift, and we characterize when a suitable acceptance rule can sustain improvement. A comparison with independently refreshed errors isolates the role of deterministic errors: two judges with the same evaluation accuracy produce opposite signs of eventual drift. A sequential test measures expected utility change with sample bounds that distinguish revision frequency, magnitude, and evaluation noise. Experiments with language model judges corroborate the theoretical predictions. We identify whether errors are deterministic as a property that determines the outcome of self-reinforcement and is not captured by evaluation accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.