Spectral Enhanced Transformer: Multi-Resolution Causal Convolution for Transformer Language Modeling
Abstract
We introduce the Spectral Enhanced Transformer (SET), which augments Transformer blocks with a learned multi-resolution causal convolutional path parallel to self-attention and causal token mixing after the MLP. FFT evaluation with zero padding preserves autoregressive causality and is numerically equivalent to direct evaluation of the same kernels. A complete 30-run WikiText-103 matrix crosses two positional encodings, five parameter-matched arms, and three paired seeds, with 61.44M training tokens per run. Full SET reduces mean validation NLL from 4.01723 to 3.88255 with learned positions and from 3.87901 to 3.84695 with RoPE, with every paired seed favoring full SET over the wider-FFN baseline. MLP-only mixing underperforms the baseline, and a fixed schedule beats depth-varying full SET in every RoPE seed, so the results do not establish a depth-varying-schedule advantage. Earlier synchronized timing measures 1.85 baseline step time, and longer baseline training wins under wall-clock matching. The supported result is a finite-budget equal-token quality–compute tradeoff, not an efficiency or scaling advantage; larger historical runs remain descriptive because of incomplete metadata and undertraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.