Vision Fluxer with Dynamic Global Receptive Field
Abstract
A token mixer should provide a global receptive field and input-dependent mixing at tractable cost: convolution, with its local, static kernel, offers neither, and self-attention offers both at a cost quadratic in the number of tokens. Frequency-domain convolution attains a global receptive field at , but its filter is fixed after training. Conditioning the filter pointwise on the input admits two constructions: the filter may be generated in frequency from the spectrum, or in space from the signal and then transformed. A linear generator cannot distinguish the two, and the choice between them has gone unexamined. Here we show that once the generator is nonlinear the two constructions are no longer interchangeable, and adopt the spatial one: Fluxer generates a full-resolution kernel as a pointwise projection of the input, so that the mixing amounts to a bilinear interaction between two nonlinear views of a single feature map. Three separations position Fluxer relative to static spectral filters, local convolution, and frequency-pointwise masks; our experiments bear most on the last: in the same backbone, a frequency-pointwise generator recovers only a small part of the gap between a static filter and Fluxer on ImageNet-1K. As a drop-in backbone, Vision Fluxer (ViF) achieves the most favorable accuracy–compute trade-off among representative single-stream hierarchical backbones on classification, detection, and segmentation, indicating that convolution, once equipped with a dynamic global kernel, constitutes a competitive foundation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.