When Does Attribution Stability Improve Explanation-Guided Learning?
Abstract
Teacher attribution maps provide spatial guidance without region annotations, but the same teacher can highlight different regions across label-preserving views. Uniform alignment gives these inconsistent targets the same weight as repeatable ones. We propose Perturbation-Consistent Gated Attribution Alignment (PCGA), which averages class-matched teacher maps into a spatial target and uses their cross-view agreement to weight student alignment. Matched comparisons against uniform and prediction-derived weights test whether agreement improves learning. On Chest CT, gated cosine alignment improves AUC over conventionally tuned uniform alignment, although matching the auxiliary-gradient strength removes that advantage. Prediction-confidence weighting leads on RADGEN. Stability weighting has the highest mean AUC on CIFAR-10, with gains that vary across seeds. Gradient measurements and spatial interventions clarify the distinction: map agreement chiefly controls alignment strength in the Chest CT comparison and does not consistently identify regions that affect predictions. Cross-view agreement can therefore set the strength of attribution supervision, but its value as an image-ranking signal depends on the task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.