THMAP : Temporal HeatMap for text-to-video Intelligence
Abstract
A typical video contains more visual evidence than vision language models (VLMs) can process effectively, making query-relevant events easy to miss. Existing approaches either expand the visual context of VLMs at high training and inference cost, or train an external model to predict query-conditioned temporal regions as a preprocessing step, often requiring task-specific supervision on large training datasets. We study whether relevance is already encoded by the VLM that interprets the selected frames, and introduce THMAP. THMAP extracts a query-conditioned temporal relevance heatmap from text-to-vision token interactions that occur within a small set of attention heads grounding the user query in visual context. We decompose the outputs of these heads in a mechanistically interpretable manner to measure token-level relevance, and aggregate these signals into a temporal relevance heatmap at the desired granularity. To sharpen this heatmap, we learn a lightweight low-rank adaptation using a differentiable ranking objective and a modest amount of labeled data, requiring one to two orders of magnitude less supervision than typical existing approaches. The same heatmap scores clips for highlight detection and selects visual context for temporal grounding and video question answering. On highlight detection, THMAP improves over the previous state of the art by relative gains of in mAP and in Hit@. For temporal grounding, it improves mIoU by upto over a strong trained selector under an eight-frame budget, while remaining competitive across two VLMs for video question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.