SeekSpan: Calibrating Temporal Extent for Video Temporal Grounding
Abstract
Video temporal grounding predicts the position (the midpoint) and the duration (extent) of an event in a video. Existing training-free verification methods replace the generated interval with one built from clip or window scores, even when the generated interval is well placed and only its duration is wrong. We propose SeekSpan, a training-free method that corrects the duration of a predicted interval around its original midpoint. A frozen video-language model, the verifier, views the full video and scores each temporal window by the yes–no log-probability margin for the queried event occurring within it, and the scores across windows estimate how long the event lasts. We take the geometric mean of this estimate and the predicted duration and place the start and end times around the original midpoint, adjusting at video endpoints when needed. With one frozen verifier shared across generator families, SeekSpan improves 12 of 13 configurations on Charades-STA, including all six thinking-mode configurations. The gains exceed one mIoU point in 11 configurations and reach 9.2 points. With a frozen Qwen3-VL-8B generator, SeekSpan reaches 57.8 mIoU on the full Charades-STA test set, above every published training-free method and above the re-measured systems trained without Charades-STA annotations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.