Neutralizing 0th-Order Residual Interferences in Layer-wise Relevance Propagation via Geometric Trajectory Realignment
Abstract
Standard Layer-wise Relevance Propagation (LRP) operates under the premise that 0th-order residual terms, even if not strictly eliminated, do not affect the essential properties of relevance propagation. However, this study reveals that when introducing a simple constant 0th-order term as a minimal 0th-order probe, the linearity of the distribution persists under constrained input methods, causing noise to accumulate during relevance propagation. This phenomenon induces a deviation () within the - parameter space and misaligns geometric trajectories, thereby structurally violating LRP's conservation principle and ultimately leading to fatal signal collapse in deep layers. To empirically validate this theoretical distortion mechanism, we introduce the -rule—a normalization mechanism that neutralizes the linear effects of the 0th-order terms by aligning geometric trajectory slopes to unity. Our analysis of CNN architectures (e.g., VGG, ResNet, DenseNet) confirms the validity of our hypothesis: the -rule maintains a constant variance of , effectively suppressing distribution-induced noise and improving explanation fidelity. Conversely, the marginal improvements observed in Vision Transformers (ViTs) expose the limitations of the simplified constant 0th-order assumption in complex non-linear models. Despite this boundary of validity, our empirical study unequivocally proves that the distribution of 0th-order terms alone generates critical noise in LRP, establishing an essential theoretical foundation for designing novel rules to address non-linear residual dynamics in future Transformer architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.