acceptodds
Under review as a conference paper at ICLR 2027

What Do Self-Improvement Methods Optimize For? An Empirical-Mechanistic Study

Abstract

Self-improvement methods such as debate, simple bootstrap, Gibbs sampling, and internal coherence maximization (ICM) can improve language model accuracy without external supervision, yet the signal they optimize remains unclear. What drives the gains, and when? We test the hypothesis that these methods improve accuracy through *coherence optimization* (Qiu et al., 2026), finding behaviors that are jointly predictable under a model's prior. For debate, bootstrap, and Gibbs sampling, we find that, overall, coherence and training-pool accuracy increase in sync, and rejecting update steps that increase coherence also lowers accuracy, suggesting a causal relation. For the remaining method, ICM, where ablation is inapplicable, Qiu et al. (2026) proved its analytical equivalence to coherence optimization under an information saturation condition, which we experimentally verified. Our results suggest that coherence optimization provides a common objective for understanding and designing feedback-free self-improvement methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.