Knowing Less, Teaching More: Weak-to-Strong Generalization from Partial Knowledge
Abstract
Weak-to-strong generalization (W2S) occurs when a strong student trained on a weak teacher’s labels outperforms the teacher on unseen samples. We study how a strong student can exploit supervision from a weak teacher that reliably recognizes an easy feature but has only partial knowledge of a hard feature, producing stochastic label noise on hard-only samples. For a two-layer ReLU network trained by gradient descent, we show that hard-feature learning approaches a saturation level determined by the feature-strength-weighted balance between correctly and incorrectly labeled hard-only samples. Overlap samples containing both features initially promote hard-feature learning, but their contribution vanishes as the easy feature is learned. Nevertheless, the student retains a positive fraction of the saturation level while fitting incorrect labels through sample-specific noise, which has little effect on unseen samples. Consequently, its test error converges to zero during a prescribed late-training interval, even after fitting every weak label. Experiments with ReLU networks and language models show that, at fixed label accuracy, assigning correct labels to samples with stronger hard features improves feature learning and student accuracy. A teacher-selection study further shows that hard-feature knowledge can be a better criterion for assessing weak-supervision quality than validation accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.