Beyond Frame-Wise Encoding: When Do Video-Native Representations Help Gait Recognition?
Abstract
Large vision models (LVMs) have advanced gait recognition, but most pipelines retain the GaitSet-style design: an image LVM encodes each frame independently before a gait head aggregates the sequence. Video LVMs instead encode clips jointly, producing video-native representations shaped by cross-frame context. We study when these representations help and what they change before the gait head. Across several benchmarks, frozen video LVMs remain competitive on within-domain evaluations and transfer better to unseen domains. To isolate the effect of clip-wise encoding, we evaluate the same V-JEPA 2.1 checkpoint frame-wise and clip-wise, using the same training samples and gait-head architecture. Clip-wise encoding has little effect on within-domain accuracy but improves cross-domain recognition, with the largest gains under nighttime capture and clothing changes. Under controlled occlusion, its advantage depends on temporal coherence rather than on whether hidden regions reappear: the advantage grows when a body region remains hidden throughout the sequence but reverses when visibility shifts randomly across frames, in which case frame-wise representations perform better. At the frozen-encoder level, clip-wise context stabilizes representations of still-visible regions and makes unchanged observations more identity-discriminative, thereby changing what the gait head pools rather than merely how it pools. Gait research has long focused on refining how features are pooled; video-native representations offer a way to improve the features themselves. Our results identify where this benefit emerges and where clip-wise context becomes a liability, positioning video-native representations as a promising foundation for robust gait recognition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.