ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
Abstract
Video-capable vision-language models score above 80% on popular benchmarks, yet they systematically fail at something more fundamental: spatial-temporal binding, i.e., associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions that decomposes this capability into five orthogonal diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Every ground-truth answer is derived deterministically from AVAv2.2's 1.58 million per-second, per-person annotations, no language model generated our labels. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Evaluating 20 VLMs reveals a sobering gap: the best open-weight model achieves 68.8% weighted average accuracy against 91.0% for humans, with gaze detection collapsing near chance despite 98.6% human accuracy. A coordinate-variant ablation, replacing visual bounding-box overlays with text coordinates, exposes up to a 19.2-point grounding deficit concentrated exclusively on diagnostics that require person-identity binding. A binding-trap analysis shows 13 of 16 models systematically select the wrong actor's action on D2. ActionLens is not another leaderboard; it is a probe: each diagnostic score is an independent measurement of a capability a model either has or does not. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276, with ActionLens integrated into lmms-eval for one-command reproduction across any supported VLM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.