acceptodds
Under review as a conference paper at ICLR 2027

Align2Refine: Diagnosing and Refining Video Representations through Spatial Latent Alignment

Abstract

Pre-trained video encoders can produce spurious tokens that fail to represent the visual content at their spatial locations. Prior work derives identification criteria from internal activations and token interactions, but extending these analyses to video is complicated by architectural diversity and complex spatiotemporal attention. We approach this challenge through task feedback: the spatial association between prediction errors and local feature fragmentation suggests that task responses can help locate spurious tokens. To obtain this feedback without downstream annotations, we formulate spatial latent alignment. We find that encoder representations capturing high-level semantics or low-level visual details can both serve as diagnostic references for this task. Building on these insights, we introduce Align2Refine, a framework for post-hoc refinement of video encoders on unlabeled videos. To make alignment feedback more discriminative for token identification, we use deliberately imbalanced learning to reinforce differences in fitting progress. The resulting token partition guides selective refinement using the original encoder's features as supervision. Across five heterogeneous video encoders, Align2Refine delivers substantial improvements in both spatial prediction and temporal correspondence under frozen-backbone evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.