Structured Temperature Scaling for Group Knowledge Distillation
Abstract
Knowledge distillation (KD) commonly uses temperature scaling to soften the teacher distribution and transfer dark knowledge. In this work, we revisit temperature from a structured prediction perspective. To this end, we derive an exact multi-group decomposition of vanilla KD into inter-group distillation (inter-KD) and probability-mass-weighted intra-group distillation (intra-KD). Inter-KD preserves the probability allocation across groups, while intra-KD transfers fine-grained conditional structures within each group. This decomposition reveals that temperature improves KD by adjusting the contribution of different probability regions, but conventional global temperature simultaneously alters inter-group allocation and flattens fine-grained intra-group structures.Based on this analysis, we propose Group Knowledge Distillation (GKD), which introduces structured temperature scaling by shifting temperature smoothing from the global category distribution to the group level. Specifically, GKD adaptively determines prediction groups based on prediction perplexity and applies temperature scaling at the group level, thereby rebalancing supervision across groups while preserving fine-grained intra-group structural knowledge. Without introducing additional grouping hyperparameters or complex structural designs, GKD only modifies the weighting strategy in the reconstructed KD formulation and achieves effective knowledge transfer.Experiments on CIFAR-100 and ImageNet demonstrate state-of-the-art Top-1 performance, providing a new perspective on temperature regulation in knowledge distillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.