Rethinking Cross-Layer Attention for Channel and Spatial Recalibration in ConvNets
Abstract
Cross-layer attention enables convolutional networks to retrieve information from earlier layers, while channel and spatial attention recalibrate features within individual blocks. We investigate how cross-layer retrieval can inform channel and spatial recalibration. We propose Cross-Layer Attention Recalibration (CLAR), a module that constructs gates for the current block by attending to channel and spatial representations accumulated within the same network stage. CLAR maintains separate memories of channel descriptors and spatial gating maps. Input-dependent queries attend to keys from preceding blocks and the current block, and the retrieved values generate channel and spatial gates that sequentially modulate the current residual features. This design preserves the local residual shortcut and avoids storing full historical feature tensors. We additionally apply stochastic history-source dropping during training to regularize cross-layer retrieval. We evaluate CLAR on CIFAR-100 and ImageNet using residual networks of different depths. In single-run ImageNet experiments trained for 100 epochs, CLAR achieves best top-1 validation accuracies of 77.49% and 78.26% with ResNet-50 and ResNet-101, respectively. On CIFAR-100, history-source dropping increases the mean final accuracy of CLAR-equipped ResNet-50 from 78.70% to 79.13% across five matched seeds. These results support cross-layer attention as a mechanism for channel and spatial feature recalibration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.