acceptodds
Under review as a conference paper at ICLR 2027

Separating Training-Data Source from Update Rate in Verified Self-Training

Abstract

Iterative self-training filters a model's own generated solutions by answer verification and fine-tunes on the survivors. Correct final answers, however, do not guarantee that repeated updates preserve performance, and lowering the learning rate to reduce forgetting also changes the data the model generates in later rounds, so retention differences cannot be attributed to the update or to the data alone. We separate these two factors for Llama-3.1-8B-Instruct over forty rounds on a fixed GSM8K stream. Generator and learner controls first establish that retention depends on the recursive update policy. A crossed design then freezes the generated training data from high-rate and low-rate runs and trains fresh receivers at each rate, with three seeds per cell. At the fixed low receiver rate, data generated under the low-rate policy improves the final-five-round near-domain accuracy change by 7.33 percentage points on average, positive in every seed. Yet every primary run still loses strictly held-out accuracy, and a reference trained on authored rationales also declines in every seed. Answer verification therefore leaves a data-source component of forgetting that update-rate tuning does not remove, and neither condition yields useful self-improvement within this protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.