TASA: TOKEN-ADAPTIVE STEERING OF ACTIVATIONS FOR VIDEO TEMPORAL GROUNDING
Abstract
Video temporal grounding (VTG) aims to identify the start and end timestamps of the video segment corresponding to a natural-language query, yet existing Video-LLMs often produce inaccurate temporal boundaries. Activation steering guides model behavior by modifying internal activations during inference. It offers a promising way to improve temporal grounding without updating model parameters. Conventional methods typically derive a single shared direction from positive-negative examples and apply it uniformly to all test samples. However, VTG involves diverse and even opposing boundary errors, making positive-negative pairs ambiguous and a single global direction insufficient for sample- and position-specific correction. We propose Token-Adaptive Steering of Activations (TASA), which uses three-stage training to distill task-loss gradients into a compact bank of steering vectors and learns a lightweight controller consisting of a token-level Router and Gate, thereby extending conventional single-direction steering to adaptive multi-directional activation steering. At inference, the controller adaptively composes the vectors and scales the intervention at each autoregressive generation position according to the current hidden state. This design combines the efficiency of shared activation steering with the flexibility of token-specific directions. Extensive experiments across multiple open-source Video-LLMs and widely used temporal grounding benchmarks show that TASA consistently delivers substantial improvements without updating the Video-LLM backbone parameters. In particular, on Charades-STA with Qwen3-VL-8B, TASA raises mIoU from 43.92 to 56.27 (+12.35) and [email protected] from 45.13 to 66.18 (+21.05). These results demonstrate that task-loss gradients can be distilled into compact, learnable activation-space interventions that effectively improve Video-LLM temporal grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.