acceptodds
Under review as a conference paper at ICLR 2027

When to Stop: Training-Free Video Temporal Grounding via Endpoint Activation Steering

Abstract

Video Temporal Grounding (VTG) aims to localize the start and end times of queried events in untrimmed videos. Despite substantial progress in event understanding, video large language models still produce unreliable timestamps. Existing improvements often rely on post-training with temporal annotations or additional inference, incurring annotation or computational costs. Contrastive activation steering offers a complementary training-free approach, but applying it to timestamp generation requires identifying the temporal components within activation differences that influence numerical boundary decisions. To address this question, we investigate how video evidence contributes to timestamp generation. We find that temporal information persists during generation, yet visual evidence provides weaker discrimination among end-time candidates than among start-time candidates. Models can recognize an event and locate its onset while still struggling to determine when it terminates. To identify the temporal evidence influencing this decision, we use endpoint-score Jacobians to locate directions that alter end-time preferences and project activation differences induced by temporal perturbations onto these directions. This construction selects components according to native endpoint sensitivity rather than activation differences alone. Aggregating the projected components across samples yields a low-dimensional subspace in late decoder layers. Causal interventions further establish that this subspace mediates the influence of temporal evidence on end-time prediction. Building on this intervention interface, we propose training-free endpoint activation steering. Using unlabeled samples, we extract a shared direction within the subspace and calibrate its relative strength offline for each model–dataset pair. At inference, we apply a single residual update immediately before the first end-time numeric token, preserving the generated start and native decoding without parameter updates or additional video forward passes. Evaluated across seven frozen backbones on TimeLens-Bench, our method improves mIoU by up to 3.62 points. Across the 4B and 8B variants of Qwen3-VL, InternVL3.5, and TimeLens2, it consistently improves both mIoU and [email protected] on all three datasets, with average gains of 2.19 and 1.46 points respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.