acceptodds
Under review as a conference paper at ICLR 2027

ActionRefiner: Refining Predictions with Inverse Verbs for Action Recognition in Egocentric Videos

Abstract

Despite advances in egocentric action recognition using vision-language models, distinguishing actions with identical nouns but inverse verbs (e.g., "open fridge" vs. "close fridge") remains challenging because these actions often only differ in their temporal order of execution. We present ActionRefiner, an approach to improve action recognition via a Causal Temporal Generation Process (CTGP) and a residual classification module. Given action intermediate spatio-temporal features, CTGP learns latent temporal representations for verbs and nouns conditioned on action context via VAE-based modelling and a domain estimator. A residual classifier then produces corrective logits that are fused with the baseline logits to predict actions. ActionRefiner consistently improves strong baselines and establishes new state-of-the-art performance on Epic-Kitchens-100 and EGTEA Gaze+. In particular, it achieves a new state-of-the-art 57.7 Top-1 accuracy on Epic-Kitchens-100 and 82.7 on EGTEA Gaze+. Notably, the majority of improvements over LaViLa on EGTEA Gaze+ is on inverse-verb actions, demonstrating the importance of learning action-aware temporal representations of verbs and nouns.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.