Knowledge Distillation for Improved Generalization in Two-Stage Classification
Abstract
Large-scale classification often uses a two-stage architecture: a lightweight candidate generator selects a small subset of classes, and a stronger target classifier scores them. Recent work has improved these systems through cross-stage training, using techniques such as knowledge distillation and joint optimization. Despite these empirical advances, the effect of the candidate-generator distribution on target-classifier training remains poorly understood. To explain this mechanism, we formulate sampled softmax training as importance sampling SGD, using the candidate generator distribution as the proposal. We also develop a variance reduction analysis of cross-stage distillation. In the context of cross-stage information asymmetry, the output distribution of the target classifier is minimax-optimal for the unbiased estimator. For the practical self-normalized estimator, alignment controls the variance and bias bounds. These results demonstrate how distillation enhances training efficiency and convergence. In addition to providing training samples, the candidate generator determines which classes are available during inference. To analyze this role, we derive conditions for teacher quality and candidate generator capacity under which positive weight distillation strictly tightens the coverage loss upper bound relative to independent training. Combining these results, we obtain an end-to-end generalization bound that considers both coverage loss and target-classifier error. Experiments on real-world datasets demonstrate that distillation yields the lowest gradient variance among non-Oracle proposals. It also improves final accuracy across three model combinations. Convergence experiments on two combinations further demonstrate faster training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.