Sparse Video Understanding Without the Heuristics
Abstract
Video recognition often requires only a handful of tokens, yet which tokens matter depends on the task. Existing approaches for sparse video computation rely on heuristics such as motion, diversity, or reconstruction to guide selection. These signals capture redundancy rather than task relevance, often retaining more tokens than necessary. We instead optimize selection directly for the task, jointly learning a lightweight selector that glimpses across space and time and a deeper extractor that processes only the selected tokens. Both are trained end-to-end through reinforcement learning, directly coupling the learning of selection with the performance of recognition. Across six action-recognition benchmarks (Kinetics-400, SSv2, Epic-Kitchens, Diving48, Jester, and Charades) and two text-conditioned tasks (A2D and CLEVRER), our approach pushes the Pareto frontier for accuracy-efficiency, outperforming state-of-the-art efficient video methods while generalizing across architectures (DINOv3, VideoMAEv2, and InternVideo2).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.