Whom Should a Student Trust? Unsupervised Reliability Estimation for Multi-Teacher Distillation
Abstract
Multi-teacher distillation provides diverse supervision for training language models. However, teacher-generated responses vary in quality and suitability for the student, making it difficult to determine which teachers to trust and which responses to learn from. Human annotations and external LLM judges can guide selection, but add cost and may introduce evaluator-specific biases, motivating unsupervised approaches. Existing student-based methods select responses using token probabilities to measure familiarity or loss gradients to characterize their effects on optimization. However, familiarity does not directly indicate learning benefit, and computing parameter gradients requires additional backward passes. Therefore, we propose Student-informed Unsupervised Reliability Estimation (SURE), which jointly estimates teacher reliability and response-selection probabilities without external quality judgments. In particular, SURE measures coherence among student gradients from different teachers’ responses to the same prompt without backpropagating through the frozen student. It then models this evidence over a teacher-prompt bipartite graph to iteratively refine teacher reliability and response-selection probabilities, integrating within-prompt coherence with cross-prompt teacher reliability. Experiments across pretrained students from different model families show consistent improvements over other unsupervised multi-teacher distillation baselines. Further analyses show higher judged quality of the selected responses and support the value of estimating teacher reliability across prompts for curating effective multi-teacher supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.