acceptodds
Under review as a conference paper at ICLR 2027

When Does Attribution Stability Improve Explanation-Guided Learning?

Abstract

Teacher attribution maps provide spatial guidance without region annotations, but the same teacher can highlight different regions across label-preserving views. Uniform alignment gives these inconsistent targets the same weight as repeatable ones. We propose Perturbation-Consistent Gated Attribution Alignment (PCGA), which averages class-matched teacher maps into a spatial target and uses their cross-view agreement to weight student alignment. Matched comparisons against uniform and prediction-derived weights test whether agreement improves learning. On Chest CT, gated cosine alignment improves AUC over conventionally tuned uniform alignment, although matching the auxiliary-gradient strength removes that advantage. Prediction-confidence weighting leads on RADGEN. Stability weighting has the highest mean AUC on CIFAR-10, with gains that vary across seeds. Gradient measurements and spatial interventions clarify the distinction: map agreement chiefly controls alignment strength in the Chest CT comparison and does not consistently identify regions that affect predictions. Cross-view agreement can therefore set the strength of attribution supervision, but its value as an image-ranking signal depends on the task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.