LatPA: A Training-Free Latent Proactive Adapter for Streaming Video-LLMs
Abstract
Streaming video assistants must decide when to act, not just what to answer. Existing proactive triggers typically obtain this decision from an additional trained module or a yes/no query to the backbone. This leaves open how to obtain a temporal decision directly from the frozen model's representations. To this end, we introduce LatPA (Latent Proactive Adapter), a training-free alternative that extracts latent signals from a frozen Video-LLM's intermediate hidden states. The key idea is to use the video's own history as the reference for the current observation. LatPA compares the question alignment of the current frame with that of an exponential memory of preceding frames and acts when the resulting score is positive. This self-referenced rule requires neither an additional learned scoring module nor yes/no generation, and uses a shared configuration across backbones and benchmarks. Experiments on StreamingBench-PO show that LatPA improves average timing-and-answer accuracy over the evaluated learned and queried triggers with Qwen2-VL backbones. Across the evaluated model families and scales, LatPA consistently outperforms the evaluated queried triggers on the average scores of StreamingBench-PO and IPIBench. These results demonstrate that a video-specific historical reference enables effective proactive timing directly from frozen latent representations. Code is available at https://anonymous.4open.science/r/LatPA-D405.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.