acceptodds
Under review as a conference paper at ICLR 2027

Aligning Prediction Residuals: Representation-Aware Data Selection for Weak-to-Strong Generalization

Abstract

Weak-to-strong generalization (W2S), where a weaker teacher supervises the training of a more capable student, offers a path to advancing frontier models beyond what costly human expert supervision can offer. Existing W2S work focuses on model-centric aspects such as adapting the training objective. Far less is known about which data to use in order to maximize the student's improvement over its teacher. We study this problem, beginning with a characterization of how bias and structured noise in weak labels limit the performance of the student. Building on this, we develop a data selection framework, **RADS**, measuring differences in predictive capability between the two models' representations. RADS prioritizes points with features the stronger model encodes but the weaker teacher misses, and training on them yields the largest performance gains. Across nine datasets, RADS improves over the strongest baseline by more than 9% on average; this holds across data budgets. We show this advantage does not come from simply selecting samples the teacher already labels correctly, but from data with features the student can additionally learn from. The effect carries over to reward modeling, where RADS outperforms naive sampling by 10% on average. Finally, we present a probing study that estimates potential W2S gains prior to training

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.