You Only Forward Once: Rethinking Video Temporal Grounding in MLLMs as Single-Pass Parallel Boundary Probability Prediction
Abstract
Video Temporal Grounding (VTG) aims to localize the video segment corresponding to a natural-language query. Existing MLLM-based methods typically generate temporal boundaries autoregressively as textual tokens, creating a mismatch with interval-level localization and incurring sequential decoding overhead. To this end, we present YOFO-VTG, which reformulates VTG as parallel boundary-distribution prediction within a single MLLM forward pass. YOFO-VTG extracts video-aligned hidden states at the visual start and end delimiters of each temporal unit and maps them through the native language modeling head to predict start- and end-boundary probability distributions. To overcome the contextual restriction of causal attention, we adopt bidirectional attention, allowing each unit to integrate information from the full video-query sequence. This design retains the native probabilistic prediction mechanism of MLLMs while directly exploiting the complete sequence of video representations. We further introduce a two-stage strategy that first learns coarse boundary distributions with cross-entropy (CE) supervision and then refines the predicted boundaries through joint CE and differentiable IoU optimization. Extensive experiments on TimeLensBench demonstrate that YOFO-VTG achieves competitive localization accuracy while enabling substantially faster inference, highlighting parallel probability distribution prediction as an efficient paradigm for MLLM-based VTG.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.