Retrieve-then-Reason: Always-On Proactive Assistance with On-Demand Reasoning
Abstract
Always-on proactive AI assistants must continuously handle streaming multimodal information, but invoking a multimodal large language model (MLLM) at every timestep is prohibitively expensive for on-device deployment. To address this challenge, we reformulate continuous proactive monitoring as a lightweight semantic retrieval problem and present PRISM (Proactive Reasoning via Semantic Monitoring), a cascaded framework that separates continuous retrieval from on-demand contextual reasoning. PRISM first identifies potentially relevant events with high recall and then filters them according to their contextual actionability, reducing false alarms and unnecessary MLLM invocations. Specifically, lightweight Visual and State Retrievers continuously retrieve instruction-relevant visual evidence and semantic states from the incoming multimodal stream leaving contextual actionability assessment to the subsequent verification stage. On-demand Verifier then evaluates whether each retrieved candidate satisfies the requested condition using the current multimodal context, recent history, and relevant semantic states before forwarding it to the MLLM-based Responder. On the PhoStream benchmark, PRISM improves performance on future-dependent proactive tasks, substantially reduces premature responses, and achieves a performance–energy Pareto improvement over direct MLLM invocation at every step. Our results show that lightweight semantic monitoring combined with contextual verification enables efficient continuous proactive assistance while reserving expensive reasoning for potentially actionable events.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.