acceptodds
Under review as a conference paper at ICLR 2027

Class-Aware Margin Optimization: Resolving KL Divergence-Classification Inconsistency in Decoupled Knowledge Distillation

Abstract

Logit-based knowledge distillation (KD) for classification offers a cost-efficient alternative to feature-based and relationbased methods. However, logit-based knowledge distillation methods fail to explicitly model the discriminative margins between target and non-target categories. This limitation primarily stems from their overemphasis on globally minimizing KL divergence during distillation , which rigidly enforces distribution-level similarity while neglecting the inconsistency between KL divergence and classification. To address this limitation, we propose a Class- Aware Margin Optimization method for Decoupled Knowledge Distillation (CAMO-DKD). CAMO-DKD enforces the margin constraint by aligning the student’s decoupled predictions with the teacher’s decoupled distribution through a margin-sensitive loss. This enhances the separation between the target class and competing non-target categories, thereby preventing boundary collapse during distillation. The proposed CAMO-DKD method can enhance any logit-based KD approach in a plugand-play manner. Additionally, we provide a detailed theoretical analysis demonstrating the generalization capabilities of CAMO-DKD. Experiments on CIFAR-100, ImageNet and COCO show consistent improvements, while outperforming state-of-the-art KD methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.