Can Layer-wise Relevance Propagation Capture Competition Among Transformer Logits?
Abstract
Layer-wise Relevance Propagation (LRP) is a widely used framework for attributing Transformer predictions to input tokens. However, existing Transformer-based LRP methods typically propagate relevance from a single output logit, which implicitly assumes that the prediction is governed by that logit in isolation. This assumption is limiting, since what determines the prediction is not the absolute magnitude of any individual logit but the relative relationships among logits. Intuitively, a factor that contributes equally and positively to all logits does not affect the final decision and therefore should not be attributed as important, a property that current algorithms generally struggle to model. We formalize this limitation from two complementary perspectives in transformers: low-coupling projection heads and highly coupled parameters in deeper layers, showing that single-logit relevance cannot faithfully reflect inter-class competition. To address this, we propose a top-k differential weighting strategy that constructs scalar targets by contrasting the largest logit with its strongest competitors, thereby explicitly incorporating competition into the relevance propagation process. Extensive perturbation experiments show that our method, by down-weighting common-mode evidence that contributes equally to competing logits, helps LRP capture the truly target-specific features that primarily influence the target logit, and consistently outperforms state-of-the-art baselines. Our code is available at https://anonymous.4open.science/r/RI-LRP-6AF0.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.