Latent Chain-of-Thought Network for Gait Recognition
Abstract
Human reasoning advances through intermediate hypotheses and iterative revision, and Chain-of-Thought (CoT) brings this process into large language models by introducing explicit natural-language intermediate steps, achieving impressive performance on complex tasks such as mathematics and code. % However, language has intrinsic limitations, including subjectivity, verbosity, and ambiguity, ultimately limiting reasoning capability. % Analogously, visual CoT (\eg, thinking with images) relies on intermediate imagery that intensifies the information bottleneck. % Accordingly, recent research has pivoted from explicit CoT to latent reasoning. % To this end, we present CoTNet that internalizes CoT in latent space beyond explicit language or imagery, emulating a human-like “think-then-speak” mechanism. % We take gait recognition as a benchmark where networks aim to identify individuals by capturing motion patterns from videos under cross-covariate conditions. % Current gait networks typically perform one-shot temporal aggregation that prematurely collapses the full gait sequence into a single gait representation, thereby blurring intermediate motion patterns. % CoTNet introduces a Chain-of-Motion-Thought (CoMT), where Motion Thought with a causal mechanism preserves informative intermediate motion patterns and Motion Chain with a mix mechanism further integrates them to construct robust motion patterns. % Extensive experiments on CASIA-B, CCPG, SUSTech1K, CCGR-MINI, Gait3D, and GREW demonstrate that CoTNet achieves state-of-the-art performance and enables deeper analysis of motion patterns.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.