Flow Equivariant Vision Transformer
Abstract
Many visual sequences exhibit a geometric flow structure, where time-parameterized Lie-group actions describe motion and elements of the associated Lie algebra generate different flows. Standard self-attention does not inherently respect this structure, potentially limiting generalization across moving reference frames with geometric symmetries. In this work, we illustrate how this problem manifests in vision transformers, causing them to process motion-related sequences in an unrelated fashion. We then introduce a provably flow-equivariant vision transformer that generalizes to motion regimes absent from training. By connecting the geometry of continuous motion to patch-based visual representations, this framework provides an inductive bias for video understanding and spatiotemporal forecasting. Empirically, we demonstrate improved performance on both simple visual motion forecasting as well as PDE forecasting for physical systems with natural flow symmetries.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.