GEET: Gated Explicit-motion Extraction Tempo-spatial Adapter
Abstract
Few-shot fine-grained (FS-FG) action recognition requires distinguishing visually similar actions that differ primarily in subtle temporal dynamics from only a few labeled examples. Existing foundation-model adaptations often rely on broad spatio-temporal updates, introducing excessive adaptation capacity for this setting. We propose **GEET** (**G**ated **E**xplicit-motion **E**xtraction **T**empo-spatial Adapter), a parameter-efficient adapter for a frozen CLIP backbone that prioritizes temporal specialization while limiting spatial adaptation. GEET separates appearance evolution from explicit frame-to-frame motion and combines them through gated temporal fusion, with lightweight spatial refinement grounding the explicit motion signal in locally coherent spatial evidence. This focused design requires only 3.8M tunable parameters and video input alone, GEET achieves state-of-the-art FS-FG recognition across diverse benchmarks with substantially lower latency and higher throughput, while remaining robust to diverse visual corruptions and temporal disruptions. Source code is attached in the suppl. material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.