acceptodds
Under review as a conference paper at ICLR 2027

ReContext: Context-Coupled Matching for Multi-Event Video-Text Retrieval

Abstract

In multi-event video–text retrieval, a description refers to selected parts of one or several activities within a video. Local actions and objects help locate relevant content and distinguish similar activities, while the surrounding activity provides context for judging their relevance to the description. Pooling the strongest local and contextual matches independently discards their pairing, making different local–activity associations indistinguishable when their pooled maxima coincide. For descriptions that span several activities, relevance also depends on combining support from different contexts. To this end, we introduce ReContext, which supplements each context's compatibility with the description using matches from its own constituent spans, then selects the highest-scoring context. We combine this context score with an aggregation of each description fragment's best match across the video, yielding a shared bidirectional score using frozen representations. Our analysis shows that coupled selection can retain distinctions between local–context configurations that independent pooling assigns identical local and contextual maxima. Experiments on Charades-Event and ActivityNet-Event show improved retrieval in both directions, with controlled comparisons examining the contributions of membership coupling and cross-activity support.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.