HMER: Hyperbolic Event Memory and Reasoning for Streaming Text-to-Video Retrieval
Abstract
Existing text-to-video retrieval methods usually perform offline cross-modal matching on complete videos, which limits their applicability to streaming scenarios where videos arrive continuously and future content is unavailable. To address this issue, we study Streaming Text-to-Video Retrieval (STVR), which requires a model to identify relevant video streams while determining whether the target event is occurring at the current moment. This setting poses three main challenges: incomplete visual evidence, interference from past events, and unstable moment-level event predictions. To tackle these challenges, we propose Hyperbolic Memory-Based Event Reasoning (HMER), which jointly models event context, cross-modal hierarchy, and moment-level event states. Within HMER, we develop a Radial Valley-Guided Event Memory (RVG-EM) module to estimate event transitions from changes in hyperbolic radial trajectories. It aggregates historical information only within the current event, enriching local visual evidence while reducing cross-event interference. Since a video segment provides only local evidence for the complete event described by the text, we model the query and video segment as a general–specific semantic relation. Based on this view, we propose the Angular Cone Separation (ACS) loss to organize matching and non-matching representations with angular cone constraints, improving hierarchical alignment and cross-modal discrimination. Because semantic relevance alone cannot determine whether the target event is currently occurring, we further propose a Streaming Query-Aware Event Reasoning (SQ-AER) module. It jointly predicts query-matching confidence and event boundaries and combines them with historical high-confidence predictions for online event assessment. This enables the model to determine when retrieval results should be output. Experiments on DiDeMo, TVR, and Charades-STA show that HMER effectively models event context, cross-modal hierarchical relations, and moment-level event states, leading to consistent improvements on the STVR task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.