acceptodds
Under review as a conference paper at ICLR 2027

Recurrent Vision in Parallel: Spectral Convolutional State Space Models

Abstract

Brains and deep neural networks share the goal of building rich and robust visual representations but pursue it through different means. Cortical circuits route sensory signals through re-entrant feedback, refining their representations iteratively over time. Deep neural networks (DNNs), by contrast, refine theirs across layers of distinct parameters, in a single feedforward pass. This distinction becomes especially pronounced for high-dimensional signals like video, where capturing rich spatiotemporal structure either demands many sequential refinement steps or a rapid growth in model size to encode it all at once. Brains take the first route via recurrence, but its sequential nature does not translate easily to GPUs. Here, we close this gap with a parallelizable state space model (SSM) formulation that lifts convolutional recurrence into the spectral domain, where composing convolutions reduces to pointwise multiplication and yields linear-time scaling. We instantiate this idea as spectral convolutional state space models (sCSSMs): a drop\-in primitive that composes with other SSM circuits including Mamba and Gated Delta Networks. A single sCSSM layer rivals human accuracy on visual reasoning challenges where transformers fail, and reproduces the time–accuracy tradeoff that humans display on these tasks. Replacing self-attention with sCSSMs on ImageNet-1K and as a LoRA adapter on a video generation model beats Standard and Transformer baselines with fewer parameters on VideoPhy-2.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.