Beyond Optimism: Selective Temporal Reinforcement Learning for Video Grounding
Abstract
Video temporal grounding localizes the temporal boundaries and occurrences of queried events. Existing methods either return a final prediction in a single pass or iteratively correct predictions based on candidate validity and overlap. The former leave residual localization errors uncorrected, while the latter overlook the benefit and degradation risk of changing the current answer. Our diagnostic study across three benchmarks shows that accepting the candidate with the highest verifier score degrades more queries than it improves. To address this issue, we propose SeTeR, a Selective Temporal Reinforcement Learning framework that learns the value of temporal edits and decides when to apply them. SeTeR revisits regions identified by a frozen grounder to propose boundary adjustments and splits of merged occurrences. We further introduce a learnable revision verifier that compares the current and proposed interval sets, jointly predicting improvement probability, degradation risk, and positive gain under supervision from the change in complete answer utility. Conditioned on these signals, a reinforcement learning policy selects edits or stops, optimizing grounding gains with penalties for degradation and unnecessary modifications. Experiments with TimeLens2-8B demonstrate improvements across all twelve reported metrics on three temporal grounding benchmarks, including mIoU gains of 1.22, 1.58, and 1.35 points. Transfer from QVHighlights to OMTG further improves event count accuracy by 8.12 points and F1 at the event level by 5.85 points, demonstrating the effectiveness of SeTeR for both boundary localization and event separation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.