acceptodds
Under review as a conference paper at ICLR 2027

Stream-of-Thought: Always-On Reasoning for Reactive and Proactive Assistance

Abstract

Live multimodal assistants, the systems that continuously stream a user's egocentric viewpoint, make persistent perception of the user's environment practical. Yet their interaction paradigm remains *reactive*: reasoning begins only once a user speaks, and the incoming stream goes unexamined until then. We propose **Stream-of-Thought** (SoT), which enables an online vision-language model to simultaneously and adaptively emit thought tokens as it perceives a streaming input, and to decide on its own when to turn those thoughts into speech. Thoughts cached across a session serve both modes of assistance: (1) *reactive*, where complex questions are answered faster and long-horizon tasks are tracked because much of the required reasoning is already done, and (2) *proactive*, where the same cached thoughts ground the help that the user never asked for. Since unconstrained “daydreaming” only lengthens context, we train thought to be *utility-directed*, optimal under both objectives: a two-stage SFT over a curated SoT-SFT dataset, followed by **Think–Speak Group Policy Optimization** (TSGPO), a novel RL algorithm that credits thinking and speaking separately. We further introduce **UnpromptedBench**, a manual collection of unprompted, open-ended assistance in egocentric streams that tests what is said and when. Through extensive evaluations, we show that a single SoT model combines strong reactive *and* proactive capabilities under a tight latency budget, outperforming prior streaming systems at comparable scale and rivaling far larger single-regime specialists.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.