ReFocus-DETR: Revisiting Saliency and Memory for Moment Retrieval and Highlight Detection
Abstract
Moment Retrieval (MR) and Highlight Detection (HD) provide complementary views of query-conditioned video relevance, yet most DETR-based systems connect them only once: an early saliency map reweights encoder memory, after which a decoder localizes moments from a fixed representation. Errors in this map therefore persist, while increasingly precise spans cannot repair the features that produced them. We introduce ReFocus-DETR, a closed-loop detector with two complementary feedback paths. Bidirectional Saliency–Span Interaction (BSSI) converts saliency peaks into parameter-free temporal proposals and writes their span semantics back to clip memory through position-Gaussian attention. Iterative Memory Refinement (IMR) then uses the shared decoder's current spans to revise saliency and update memory along position, channel, and intra-moment coherence axes. Near-identity gates and stop-gradient span conditioning keep the recurrent updates stable. Across QVHighlights, TACoS, Charades-STA, and TVSum, ReFocus-DETR consistently improves both localization and saliency metrics. The gains persist across feature regimes and exceed parameter-matched and training-only controls, while adding 2.6% parameters to the detector.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.