acceptodds
Under review as a conference paper at ICLR 2027

Decoupling Temporal Action Ownership and Actor Geometry for Spatio-Temporal Action Localization

Abstract

Spatio-temporal action localization must identify what action occurs, when it starts and ends, and where each actor is in every frame. We introduce YOLO-ST, a detector that separates temporal action ownership from spatial actor geometry. A Kinetics-pretrained VideoMAE-L pathway supplies semantics, a frozen (2+1)D pathway that sees every frame supplies motion near likely actors, and a frozen COCO-pretrained YOLO11-L pyramid supplies frame appearance. Dense per-frame heads own geometry, tube queries own actor identity and temporal extent, and at inference a training-free snap moves each linked tube box toward the dense detection of the same class. Evaluating on UCF101-24, we find that published frame mAP is not one metric: on identical predictions, the conventions in use span 6.6 points, from 88.47 under the all-frame evaluator that ROAD recommends to 95.10 under our re-implementation of the one released with YOWOFormer. Under the latter our checkpoint exceeds YOWOFormer-L's reported 93.35; under ROAD's two all-frame evaluators, whose protocol BMViT cites, it scores 88.47 and 89.30, below BMViT's 90.7, because it fires on frames outside the annotated actions, a failure that gating by the tubes' own temporal extent does not repair. Video mAP is 89.82 / 73.39 / 34.97 at spatio-temporal overlap 0.2, 0.5, and 0.5:0.95, level with STAR-L on the first two and 0.83 below it on the third. Substantial gaps to the strongest results remain on JHMDB-21 (91.5 frame), AVA v2.2 (33.23) and MultiSports (42.11 frame).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.