Prediction Distillation Can Break Editability in Concept-Embedding Models
Abstract
Concept models expose human named attributes for users to correct, but these edits help only if the classifier acts on them. We ask whether matching a teacher's predictions through knowledge distillation preserves this ability in concept embedding models. Edits replace concept probabilities but retain image dependent embeddings, allowing a student to preserve teacher distinctions absent from the annotated concepts. On CUB with frozen ResNet 18 features, distillation raises task accuracy from 59.3% to 64.9% but lowers accuracy after full concept correction from 99.9% to 88.0%. Predictions also show much weaker alignment with supplied alternative class concepts. With frozen features, retaining distillation while shuffling teacher targets among images with identical concepts or replacing them with mean teacher logits largely restores correction benefits. These controls respectively break image and target pairing and remove image specific variation within each group. The pattern extends across image encoders, to a second dataset, and to a nonlinear classifier. Under end to end training, the effect is weaker and mean logit targets provide only partial recovery. Changes in correction accuracy are small on datasets with incomplete concepts. Prediction after an edit should therefore be evaluated alongside ordinary accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.