Towards Training-free Proactive Responses for Video Multimodal Large Language Models via Cumulative Information Estimation
Abstract
Despite great success, existing multimodal large language models (MLLMs) still look “dumb” in online video interaction, primarily due to the lack of proactive response capability. In this paper, we propose a novel and plug-and-play approach for MLLMs, termed Cumulative Information based Active response (CIAct). Unlike existing methods that typically frame proactive response as a multi-class prediction of model states, CIAct formulates it as a problem of cumulative information estimation, i.e., allowing the MLLM to respond to questions once relevant information is accumulated enough. This definition not only allows CIAct to leverage existing lightweight vision-language models to help perceive queryrelated information, but also avoids the expensive MLLM-based video watching and dedicated SFT tuning. Moreover, CIAct is also equipped with a carefully designed memory replay scheme, helping it to cope with different online QA tasks efficiently, such as backward tracing, real-time perception, and forward active responding. To validate CIAct, we apply it to two representative offline MLLMs, and compare to a set of proactive response models and methods on both online and offline benchmarks. The experimental results not only show its superior performance than existing methods on streaming benchmarks, e.g., +4.04% on OVOBench, but also confirm its great benefits to existing offline MLLMs. For instance, CIAct helps Qwen3-VL save about 38% of watching time on LVBench, while obtaining a +2.4% performance gain. Our code is in the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.