acceptodds
Under review as a conference paper at ICLR 2027

Towards Training-free Proactive Responses for Video Multimodal Large Language Models via Cumulative Information Estimation

Abstract

Despite great success, existing multimodal large language models (MLLMs) still look “dumb” in online video interaction, primarily due to the lack of proactive response capability. In this paper, we propose a novel and plug-and-play approach for MLLMs, termed Cumulative Information based Active response (CIAct). Unlike existing methods that typically frame proactive response as a multi-class prediction of model states, CIAct formulates it as a problem of cumulative information estimation, i.e., allowing the MLLM to respond to questions once relevant information is accumulated enough. This definition not only allows CIAct to leverage existing lightweight vision-language models to help perceive queryrelated information, but also avoids the expensive MLLM-based video watching and dedicated SFT tuning. Moreover, CIAct is also equipped with a carefully designed memory replay scheme, helping it to cope with different online QA tasks efficiently, such as backward tracing, real-time perception, and forward active responding. To validate CIAct, we apply it to two representative offline MLLMs, and compare to a set of proactive response models and methods on both online and offline benchmarks. The experimental results not only show its superior performance than existing methods on streaming benchmarks, e.g., +4.04% on OVOBench, but also confirm its great benefits to existing offline MLLMs. For instance, CIAct helps Qwen3-VL save about 38% of watching time on LVBench, while obtaining a +2.4% performance gain. Our code is in the supplementary materials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.