ProV-Mem: Proactive Memory for Streaming Video Question Answering
Abstract
Proactive streaming video question answering requires a model to continuously monitor an incoming stream and autonomously decide when to speak, what to say, and whether to emit a response. Existing streaming systems face a fundamental difficulty: retaining long visual histories is costly, whereas discarding them can lead to missed events and repeated responses. We propose **ProV-Mem**, a training-free proactive memory framework built around a frozen VLM backbone. ProV-Mem decomposes the proactive decision into three lightweight memory tiers: a temporal-trigger memory controls when to query the backbone and recovers from prolonged silence, an instructional memory guides what to say through a fixed prompting prior for concise and visually grounded generation, and an episodic memory determines whether to emit by filtering near-duplicate responses. Under an idealized conditional-sufficiency assumption, we derive a stagewise approximation bound relating the decomposed process to a full-state joint policy. We further introduce Cost-Aware Area Under the Curve (**CAUC**), which extends PAUC by explicitly discounting redundant, overly frequent, and unnecessarily long emissions. Experiments on the four scenarios of ProactiveVideoQA show that ProV-Mem achieves the best CAUC across all scenarios and the best PAUC on WEB, EGO, and TV, while maintaining low per-response inference latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.