acceptodds
Under review as a conference paper at ICLR 2027

Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

Abstract

Understanding facial expressions in unconstrained videos requires modeling subtle and localized facial dynamics. Despite substantial advances enabled by large-scale self-supervised pretraining, recent Vision Transformer (ViT)-based video models often remain biased toward dominant global motions and coarse temporal dynamics, limiting their sensitivity to fine-grained facial variations. To address this limitation, we propose MiRA (Marginal-induced Attention Redistribution), a plug-in framework for ViT backbones that enhances spatio-temporal selectivity for fine-grained facial cues without introducing additional trainable parameters. In its principled exact mode, it derives frame-level confidence and intra-frame concentration from post-softmax attention, capturing each frame's aggregate attention mass and its spatial concentration, respectively. Together, these quantities form a framewise importance score, which is then transformed into a prior for reallocating attention mass across frames based on global salience and within-frame localization. To further improve efficiency, we introduce flashLite mode, a lightweight pre-softmax approximation that derives framewise logit modulation from key-based proxy statistics and incorporates it into FlashAttention kernels while preserving the effectiveness of the exact formulation. Extensive experiments demonstrate that MiRA achieves state-of-the-art performance among video-only methods while remaining competitive with multimodal approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.