acceptodds
Under review as a conference paper at ICLR 2027

What Does Logit Distillation Preserve in Pre-Trained Class-Incremental Learning?

Abstract

Distilling stored predictions is a standard remedy for forgetting in class-incremental learning. However, in stored-logit replay, many of the stored targets describe classes the model has never trained on: with a fixed classification head, every saved example carries outputs for classes not yet observed. Discarding these unsupervised constraints and distilling only over observed classes appears to be a harmless simplification. In this paper, we investigate what these constraints preserve and show that the simplification is not harmless. Removing these constraints consistently helps the model acquire new classes, but its net effect on accuracy reverses between training settings. Amplifying the reduced distillation signal to the strength of the full constraint does not recover full-distillation accuracy: the remaining loss concentrates on discrimination between classes learned at different times rather than within any single task. Probing the removed terms reveals why. The apparently unsupervised coordinates act strongly on classes learned after the targets were stored, and thereby restrict updates that later tasks depend on. Stored-logit constraints thus preserve cross-task discrimination in a way that their magnitude alone does not explain, and their value must be judged by which updates they restrict rather than by how strongly they constrain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.