A Dual-memory Framework for Language-driven Action Localization in Streaming Videos
Abstract
Language-driven action localization in streaming video aims to temporally ground natural language descriptions in an ongoing video stream, using only past and current frames. In contrast to offline settings with full video access, the streaming scenario lacks future context, leading to substantial uncertainty when recognizing partially observed actions. To address this issue, we draw inspiration from human cognition, where short-term experiences are gradually abstracted and integrated into long-term knowledge that helps interpret or anticipate actions under limited observation. We propose a dual-memory framework, which consists of two complementary memory modules: (1) an intra-memory module that encodes short-term temporal context from historical frames within the current video stream, capturing scene-specific static cues and scene-agnostic action cues, and (2) an inter-memory module that maintains long-term prototypical action representations distilled from previously observed videos. Importantly, the inter-memory is formed by integrating scene-agnostic action representations from intra-memory instances across multiple videos, reflecting how humans consolidate long-term knowledge from repeated short-term observations. During localization, the intra-memory provides the scene-specific reference, while the inter-memory retrieves plausible scene-agnostic action trends. By fusing them, our model synthesizes scene-specific future representations, compensating for the absence of future frames. Comprehensive experiments on several benchmark datasets demonstrate that our method consistently outperforms existing methods under the streaming setting, validating the effectiveness of our cognitively inspired dual-memory design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.