From Weak Teachers to Biased Students: Bias Propagation in Weak-to-Strong Generalization
Abstract
Weak-to-strong generalization (W2SG) trains a capable student using supervision from a weaker model, but most evaluations ask only whether the student recovers task capability. We instead study whether a systematic behavioral difference in the teacher is associated with a downstream student difference. In a controlled Qwen2.5 7B-to-32B setup, we fine-tune biased and clean teacher variants and compare students trained on biased and clean pseudo-labels across factual error under incorrect user assertions, persona-conditioned opinion agreement, and pressure-induced answer changes. We introduce the Bias Amplification Factor (BAF), the ratio between the two student target-behavior rates, and report it together with absolute rates and signed teacher and student contrasts. Across the evaluated checkpoints, the factual-assertion and pressure endpoints have BAFs of 2.34 and 2.84, whereas the opinion endpoint has a BAF of 0.87. A contrast audit separates three regimes: the factual student gap expands relative to the teacher gap, the opinion gap reverses sign, and the pressure gap remains positive but contracts. For factual behavior, an exact-count partition further distinguishes joint same-error, joint different-error, teacher-only, and student-only outcomes. The results show that supervision-linked behavioral changes can differ qualitatively across endpoints and motivate reporting behavioral contrasts alongside capability-oriented evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.