acceptodds
Under review as a conference paper at ICLR 2027

The Limits of Imitation: A Theory of Recalibration in Weak-to-Strong Generalization

Abstract

A student trained only on a teacher's predictions can become more accurate than the teacher, although imitation adds no information. Relative to a prior on the target, every predictor's risk splits into ignorance, the risk of the best readout of what it reveals, and calibration error, its distance from that readout. Imitation on fresh inputs cannot reduce ignorance, whatever the downstream algorithm; this one bound contains both the teacher's unobserved signal and the Bayes limit of its labels. A student with no knowledge of its own can gain only by recalibrating its teacher. In a linear Gaussian model with a sketched teacher, fitting pseudo-labels applies a spectral filter to the teacher and pays a leak: the energy the filter removes returns as noise. In this calculus, a minimum-norm teacher is overconfident exactly in the observed directions whose signal variance falls below one threshold, and imitating slightly less than a clone helps, to first order, when that overconfidence exceeds an explicit leak rate. With high-variance signal directions in dimensions, the limiting optimal budget of a minimum-norm student equalizes the fraction of the teacher's error it keeps and the fraction of the removed error it leaks. Along a chain of identical stages, energy and knowledge decay on two clocks, with an underconfident phase between them; at equal total budget, two stages can beat every single-stage ridge. In power-law formulas, interpolating pipelines acquire knowledge at the Bayes exponent, yet their risk plateaus at pure calibration error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.