acceptodds
Under review as a conference paper at ICLR 2027

SpaTAD: Preserving Spatial Structure in Temporal Action Detection

Abstract

Temporal action detection (TAD) requires modeling how visual evidence evolves over time. Most detectors, however, pool each frame or snippet into a single fea- ture vector before temporal modeling. This is effective when action evidence can be summarized globally, but can become restrictive when an action is defined by where evidence occurs and how it moves, such as the path a vehicle takes through an intersection, or when concurrent actions rely on different regions of the frame. We therefore ask how far into a temporal action detector explicit spatial structure should be kept. We introduce SpaTAD, a TAD model that keeps spatial tokens separate until action scores are computed. Rather than pooling each frame into a single vector, SpaTAD maintains a causal, location-bound recurrent state for each spatial token, predicts action scores at individual tokens, and only then aggregates them across space into frame-level action scores. On MultiTHUMOS, Charades, ATARS, and ROAD-Waymo, SpaTAD improves over the strongest matched base- line by an average of 11.6 Seg-mAP and 6.5 Det-mAP. Pooling before class pre- diction hurts performance on all four benchmarks, while collapsing spatial struc- ture before temporal modeling causes an additional 5.07 Det-mAP and 10.41 Seg- mAP drop on ATARS. These results suggest that dense temporal modeling alone is not always sufficient: preserving spatial organization through class prediction can be beneficial, and location-bound temporal memory appears especially useful when actions depend on how evidence moves over time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.