FinishStrong: Learning to End Well after Imperfect Start in Video Temporal Grounding
Abstract
A poor start lowers the ceiling, but does not negate the value of finishing well. In autoregressive video temporal grounding, models systematically exhibit an endpoint performance asymmetry, predicting ends significantly worse than starts. A committed flawed start acts as an immutable physical constraint, irreversibly capping the attainable overlap as a sunk loss. While models should minimize the remaining conditional regret, standard post-training paradigms fail to teach this adaptation. Specifically, teacher-forced SFT suffers from strict prefix exposure bias, and holistic trajectory evaluation in standard GRPO inadvertently penalizes conditionally optimal ends if they follow a poor start. To address this, we introduce \method, a novel two-stage post-training framework explicitly designed to teach optimal continuations. First, Online-start SFT exposes the end prediction to the model's own generated contexts. Second, Conditional-tree GRPO restructures the sampling topology, centering end rewards exclusively within the same-start branch to mathematically isolate the conditional end quality from the inherited start penalty. \method-9B achieves a state-of-the-art 61.7% overall IoU on TimeLens-Bench, remarkably outperforming flagship closed-source models like Gemini-3.8-Flash and Qwen3.8-Max. Crucially, when forced to continue from imperfect starts, it reduces the subsequent end prediction error by 36.5%, proving its exceptional ability to end well after flawed start. Code will be open.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.