ReContext: Context-Coupled Matching for Multi-Event Video-Text Retrieval
Abstract
In multi-event video–text retrieval, a description refers to selected parts of one or several activities within a video. Local actions and objects help locate relevant content and distinguish similar activities, while the surrounding activity provides context for judging their relevance to the description. Pooling the strongest local and contextual matches independently discards their pairing, making different local–activity associations indistinguishable when their pooled maxima coincide. For descriptions that span several activities, relevance also depends on combining support from different contexts. To this end, we introduce ReContext, which supplements each context's compatibility with the description using matches from its own constituent spans, then selects the highest-scoring context. We combine this context score with an aggregation of each description fragment's best match across the video, yielding a shared bidirectional score using frozen representations. Our analysis shows that coupled selection can retain distinctions between local–context configurations that independent pooling assigns identical local and contextual maxima. Experiments on Charades-Event and ActivityNet-Event show improved retrieval in both directions, with controlled comparisons examining the contributions of membership coupling and cross-activity support.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.