acceptodds
Under review as a conference paper at ICLR 2027

: Parallel Kinematic Selective State Space Scanners for Efficient Video Understanding

Abstract

Traditional video models relying on dense spatiotemporal attention suffer from quadratic computational costs. To circumvent these costs, recent approaches adapt image models for videos via Parameter-Efficient Fine-Tuning (PEFT) methods such as adapters. However, deeply inserting these modules incurs prohibitive activation memory overhead during back-propagation. Video state-space models (SSMs) with global scanning provide linear complexity but flatten video tokens into a one-dimensional sequence, mixing spatial and temporal ordering. To address these limitations, we propose Parallel Kinematic Selective State Space Scanners (). We retain a 2D vision backbone for spatial semantics and insert a single plug-and-play module with linear-complexity temporal scanning, bypassing the need for temporal attention or multi-layer adapters. We first explicitly extract kinematic priors via a Kinematic Prior Encoder, which captures inter-frame correspondences and feature changes. These priors subsequently drive SSMs to model temporal dependencies, adaptively modulating the update speeds and read-write strategies based on the input content at each time step. Instead of global scanning, we deploy parallel scanners along the temporal dimension for each spatial location, preserving spatial structures while limiting computational overhead. Extensive experiments on spatial-heavy and temporal-heavy action recognition benchmarks demonstrate that reaches 74.3% Top-1 accuracy on SSV2 and 84.1% on K400. The ViT-B/16 configuration fine-tunes for epochs and uses approximately less video-domain training compute than the compared VideoMamba-M configuration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.