When the Evaluator Changes: Measuring False Progress in Self-Improving Language Models
Abstract
Self-improving language models rely on evaluators both to guide learning and to measure progress. When evaluators evolve alongside the models they assess, rising scores may reflect improved behavior, changes in evaluation criteria, or adaptation to evaluator weaknesses. We study false progress: apparent improvement that does not persist when model versions are compared under a diagnostic evaluation held fixed across training. We introduce a retrospective measurement framework based on a cross-version evaluation matrix, in which outputs from successive model versions are evaluated under multiple evaluator versions. This separates changes in model behavior from changes induced by evaluator revision, while a progress decomposition quantifies how much apparent improvement survives under a common reference. To connect measurement failures to training history, we build on **Loom**, a replayable trajectory dataflow that records which supervision signals were consumed by which policy updates. This lets us distinguish three properties that are often conflated: whether measured progress remains valid under evaluator revision, whether a model state descends from disputed supervision, and whether that state remains recoverable after the evaluator is repaired. In controlled reinforcement-learning post-training experiments, we observe cases where reported scores increase while performance under a fixed diagnostic reference deteriorates. We further find that descent from affected supervision does not by itself imply irreversible behavioral failure: some intermediate states remain recoverable even after subsequent training substantially degrades reference performance. These results motivate treating progress measurement, training ancestry, and recovery as distinct questions. More broadly, they suggest that self-improving systems with evolving evaluators should be assessed through versioned, retrospective evaluation rather than contemporaneous scores alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.