acceptodds
Under review as a conference paper at ICLR 2027

C2I-CasITI-VTG: Confidence-to-IoU Cascaded Inference-Time Intervention for Video Temporal Grounding

Abstract

Video temporal grounding (VTG) aims to localize the temporal segment in an untrimmed video that corresponds to a natural-language query. Recent inference-time intervention (ITI) methods improve model behavior by steering internal activations. A natural VTG strategy is to subtract the mean activation of low-IoU samples from that of high-IoU samples, encouraging better-localized predictions. However, high IoU does not necessarily indicate that the underlying activation reliably represents successful localization. IoU and model confidence are sometimes misaligned: predictions can have high IoU with low confidence, or high confidence with poor overlap with the ground truth. Thus, IoU-only contrasts may mix inconsistent internal behaviors and produce noisy steering directions. We propose a confidence-to-IoU cascaded ITI framework for VTG (C2I-CasITI-VTG). First, confidence-based ITI calibrates the model's confidence behavior; IoU-based ITI then builds on these confidence-aligned representations to improve localization. Technically, the interventions cannot be naively combined as independently estimated steering vectors may point in incompatible directions and their superposition may cancel useful components. We therefore apply confidence-based ITI during prefill stage to reshape the subsequent decoding activations toward a confidence-conditioned representation space, from which we extract the IoU steering direction. In summary, our key contribution is to demonstrate that ITI can be cascaded to steer model behavior sequentially along complementary directions. Experiments on Charades-STA, ActivityNet-Captions, and QVHighlights show consistent gains across frozen MLLMs. On Charades-STA, our method improves Qwen3.5-9B from 53.93% to 57.47% mIoU and from 32.72% to 41.16% [email protected].

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.