Data-Size-Dependent Supervision in Weak-to-Strong Learning
Abstract
Training high-capacity models to generalize from limited data remains challenging. Teacher-generated supervision can help, but typically relies on a highly capable expert teacher. We study data-efficient weak-to-strong (W2S) generalization, asking whether pretrained weak teachers can effectively supervise stronger students with limited training data. Our experiments show that students can outperform their weak teachers at small data budgets, although their performance can saturate early as more data are added under a fixed teacher. Comparing supervision sources reveals a data-size-dependent crossover: among the sources evaluated, weaker teachers yield the best student performance at the smallest data budgets, while stronger teachers and eventually hard labels become preferable as the data budget increases. To examine this behavior, we characterize asymptotic risks in a two-stage linear regression model. Numerical evaluation of these risks reveals how the balance between bias and variance shifts with student data size, explaining changes in the preferred teacher within this model. Together, these findings highlight the potential of weak teachers for data-efficient learning and show that effective supervision depends jointly on teacher capability and the student’s data budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.