Admission Control for Training-Free Zero-Shot Temporal Action Localization
Abstract
Zero-shot temporal action localization is usually benchmarked with detectors fine-tuned on in-domain segments, and those gains do not survive a change of domain. We work out of domain throughout: a frozen image–text encoder is applied directly to test videos, and no weight is ever updated. On that encoder we build a chain of four components. A cross-video memory collects per-class evidence across the test stream, gated by confidence, and blends it into the frame relevance curve. A class-conditioned admission rule, built from a per-class text contrast direction, decides which frames the memory accepts. A video-level reliability term reads the same contrast at video granularity and rescales a video's scores against the rest of the stream. A decode threshold then places extents at a quantile of each video's own relevance curve. The resulting method improves over the state of the art on THUMOS14 and ActivityNet, at both seen/unseen class splits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.