Actions Speak Louder Than Frames: Quantifying Static and Temporal Reliance In Activity Recognition
Abstract
Video action recognition models can base their predictions on two qualitatively different sources of evidence: the visual appearance present within individual frames and the temporal dynamics that unfold across frames. Existing explanation methods typically conflate these sources within a single spatiotemporal attribution, making it difficult to determine and quantify how much of the prediction is supported by static appearance, motion, or both. To distinguish genuine temporal evidence from confidence already recoverable from static frames, we introduce Motion Headroom, which also provides a criterion for when temporal attribution is identifiable. We then introduce Reliance Probes that measure how removing appearance or temporal variation affects noun and verb predictions, and use these probes to define a Static-Bias Index (SBI) for auditing and quantifying static reliance in trained video recognition models. To localize this reliance, we use FAM (Factored Appearance and Motion), a post-hoc framework that independently optimizes appearance and motion explanations.Across EGTEA Gaze+ and Something-Something V2, the audit reveals greater static reliance on EGTEA and stronger temporal reliance on Something-Something V2 across multiple architectures. We further define an explanation-conditioned SBI that measures how much of the model's temporal reliance is captured by the temporal evidence selected by FAM. Finally, we introduce an axis-specific evaluation protocol that evaluates appearance and motion explanations against the prediction component each axis is intended to explain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.