Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding
Abstract
Understanding facial expressions in unconstrained videos requires modeling subtle and localized facial dynamics. Despite substantial advances enabled by large-scale self-supervised pretraining, recent Vision Transformer (ViT)-based video models often remain biased toward dominant global motions and coarse temporal dynamics, limiting their sensitivity to fine-grained facial variations. To address this limitation, we propose MiRA (Marginal-induced Attention Redistribution), a plug-in framework for ViT backbones that enhances spatio-temporal selectivity for fine-grained facial cues without introducing additional trainable parameters. In its principled exact mode, it derives frame-level confidence and intra-frame concentration from post-softmax attention, capturing each frame's aggregate attention mass and its spatial concentration, respectively. Together, these quantities form a framewise importance score, which is then transformed into a prior for reallocating attention mass across frames based on global salience and within-frame localization. To further improve efficiency, we introduce flashLite mode, a lightweight pre-softmax approximation that derives framewise logit modulation from key-based proxy statistics and incorporates it into FlashAttention kernels while preserving the effectiveness of the exact formulation. Extensive experiments demonstrate that MiRA achieves state-of-the-art performance among video-only methods while remaining competitive with multimodal approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.