AnchorMAE: Bridging Temporal Correspondence and Task Semantics through Frame-Level Semantic Transfer
Abstract
Ego–exo video learning must recognize activities across changing viewpoints without losing the temporal detail needed to distinguish action phases. Masked cross-view reconstruction provides dense temporal supervision, while clip-level language alignment alone does not directly specify task supervision for individual deployable visual frames. Prior methods also combine language and temporal prediction; the question examined here is how to couple dense visual reconstruction with same-video frame-level task supervision without requiring language at inference in the unpaired ego–exo setting. We introduce AnchorMAE, a shared-encoder dual-path framework with an unconditioned visual pathway and a training-only task-conditioned pathway. Dual-stage Semantic Anchor Adapters condition the latter; task-aware clip alignment and detached, same-video frame-level transfer supervise the former while masked self-/cross-view prediction retains pre-semantic visual targets. Only the visual pathway is used at inference. On AE2, AnchorMAE reports 0.8687 Phase F1, 0.8264 frame mAP@10, and 0.9966 Kendall's . Evaluations on AE2 and Ego-Exo4D, reported ablations and semantic/temporal diagnostics examine whether the proposed coupling supports both phase-sensitive correspondence and activity discrimination, while revealing remaining rank-dependent retrieval trade-offs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.