acceptodds
Under review as a conference paper at ICLR 2027

DualStreamBench: A Benchmark for Dual-System Streaming Video Understanding

Abstract

A streaming video assistant must decide when a question warrants more computation. We propose DUALSTREAMBENCH, a benchmark that evaluates this decision through paired fast and slow answers to the same timestamped query. It covers 5,619 questions over 1,392 videos, 2,190 new human-authored open-answer questions, and audited evidence-demand labels across diverse streaming tasks. Our central finding is that paired diagnosis changes resource and routing decisions that aggregate accuracy alone cannot support. Across closed- and open-source systems, including open models up to 397B parameters, slow paths overturn correct fast answers on 3.7–14.5% of questions, including 6.7–13.1% in the six full-core configurations where they raise average accuracy. Training on paired outcomes improves allocation: on a 35B model, a rescue router beats random selection given twice as many slow calls and comes within 0.02 points of always escalating with 25% fewer, whereas a frozen visual gate stays near random. On the same checkpoint, a longer visual window adds 6.57 points while 40.8× more thinking tokens add none, and complete-stream replay attributes 91.6% of routed service cost to memory maintenance. We believe DUALSTREAMBENCH makes beneficial escalation, transferable allocation, and state maintenance measurable targets for streaming video systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.