NativeQuery: Bringing Spoken Queries into Continuous Audio-Visual Streams
Abstract
Streaming omnimodal models can continuously perceive audio and visual inputs, but user queries are still typically provided separately as explicit inputs rather than arising within the stream itself. In real-world interactions, queries naturally occur as part of the continuous audio-visual stream, requiring models to recognize and interpret them while continuously processing the surrounding multimodal context. To address this challenge, we introduce NativeQuery, consisting of NativeQuery-Train, a large-scale training corpus, and NativeQuery-Bench, an evaluation benchmark for spoken-query interaction in continuous audio-visual streams. NativeQuery-Train contains 200K user-query turns, while NativeQuery-Bench contains 3,419 evaluation user-query turns across five task families and fifteen subtasks, covering grounded multimodal reasoning, conflict and ambiguity handling, persistent task execution, dynamic task management, and situated interaction reasoning. NativeQuery integrates queries into the original streams while preserving their native speakers and surrounding audio-visual context. We further introduce NativeStream, a streaming omnimodal model trained on NativeQuery-Train for such in-stream interaction. Experiments on NativeQuery-Bench reveal a substantial gap between explicit text queries and in-stream spoken queries, with Qwen 3.8 Omni Flash Realtime dropping from 64.5% to 42.8% in content accuracy. NativeStream substantially mitigates this problem, with NativeStream-7B achieving 36.2% joint accuracy on NativeQuery-Bench and 0.419 All-Global on OmniInteract, outperforming existing streaming omnimodal models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.