CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
Abstract
Autoregressive large vision–language models (LVLMs) process video by projecting video features into the LLM's embedding space as continuous visual token embeddings. Where temporal evidence is represented along this pathway, and how it causally shapes decoding, remains unclear. We present **CircuitProbe**, a circuit-level analysis framework that dissects the end-to-end video-language pathway in two stages. *Visual Auditing* localizes object semantics within the projected video-token sequence and establishes their causal necessity through targeted ablations and controlled substitutions. *Semantic Tracing* applies logit-lens probing to track the layer-wise emergence of object and temporal concepts, and uses temporal frame interventions to measure sensitivity to temporal structure. The analysis locates a consolidation interval in which visual content becomes language-aligned, together with a set of temporally specialized attention heads. Amplifying these heads in a single window fixed inside the consolidation interval (Layers 25–30), without any parameter update, improves all seven evaluated benchmarks, with significant held-out gains of +4.0% absolute on EgoSchema, +2.6% on VideoMME, and +2.4% on NExTQA, and smaller gains on the development benchmark TempCompass. These results provide interventional evidence that the identified layer interval and temporal-routing heads are functionally relevant to LVLM performance on temporal video benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.