InvAct: Reducing Scene Bias in View-Invariant Atomic Action Learning
Abstract
View-invariant action learning methods aim to match actions from videos across ego-exo views. However, because these views also share the same environment, models can match scenes rather than actions, limiting fine-grained atomic action understanding. We address this with InvAct, an atomic action embedding model designed for viewpoint and scene invariance. To reduce reliance on scene cues, we propose structured token commitment, a hard row-wise attention mask that routes each video token to a single special-token slot. Action-relevant tokens align with the class token, while non-action tokens align with register tokens that are discarded at inference. We further introduce scene-aware contrastive objective that attracts same-action clips across scenes while separating different actions within a scene. On Ego-Exo4D, InvAct improves cross-scene and cross-view atomic retrieval while reducing same-scene bias. Importantly, despite training only on atomic actions, combining its embeddings provide strong longer action retrieval without retraining. The same checkpoint transfers to unseen datasets without finetuning and provides useful latent targets for robot-policy pretraining. These results highlight scene dependence as an important consideration in view-invariant action learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.