On High-Frequency Collapse in Spectral Vision Transformers
Abstract
Modern vision architectures often emphasize low-frequency representations as depth increases. In spectral token mixers, this tendency can suppress high-frequency (HF) information—edges, textures, and fine local structure—that is discriminative for detail-dependent recognition. We study this architecture-dependent failure mode, which we call high-frequency collapse, and introduce HF-Skip: a frequency-selective pathway that applies the learned spectral gate to low-frequency coefficients while routing the upper-frequency band around it. Controlled band-filtered datasets isolate frequency content from dataset size, labels, and image identity. On ImageNet-100 at split radius r = 0.12, HF-Skip improves the HF-only arm by +12.21 ± 0.25 points but the matched LF-only arm by only +0.67 ± 0.70. Parameter-matched controls on DTD further show that HF-Skip gains +8.90 points, compared with +6.78 for a full-spectrum bypass, +4.15 for a random-bin bypass, +0.53 for an inverted frequency mask, and +0.00 for a generic spatial residual. On standard benchmarks, HF-Skip improves SpectFormer by +1.27 points on ImageNet-100 and +8.83 on DTD, generalizes to GFNet, and adds only 0.009% FLOPs. Full ImageNet-1K accuracy is tied (79.674 vs. 79.678), delimiting the benefit to regimes where fine detail matters rather than claiming a universal gain. Together, these results identify HF collapse as a task-dependent bottleneck and show that frequency-selective routing preserves useful fine detail with negligible computational overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.