acceptodds
Under review as a conference paper at ICLR 2027

DIVE: Benchmarking Full-Duplex Interaction with Visual Engagement in Continuous Streams

Abstract

Recent advances in multimodal large language models (MLLMs) have substan- tially improved spoken interaction and video understanding. Existing online and streaming benchmarks primarily assess causal multimodal perception through timestamped question answering, whereas real-world interaction requires mod- els to continuously identify task-relevant multimodal cues and decide whether, when, and how to respond. This leaves open how well current models can sus- tain real-time interaction throughout an evolving multimodal stream. To address this gap, we introduce DIVE, a benchmark for evaluating full-duplex interaction with visual engagement under continuous and realistic multimodal streams. Our benchmark organizes six tasks into three complementary interaction scenarios. (1) Proactive Intervention evaluates whether models can identify emerging vi- sual or semantic cues and intervene at the appropriate moment. (2) Continuous Engagement evaluates sustained participation through continuous visual tracking and visually grounded query interaction. (3) Interaction Coordination evaluates how models coordinate listening and speaking under overlapping and multi-party interactions. DIVE contains 525 audio/video samples with fine-grained human annotations and precise temporal metadata, covering six task-oriented interaction settings grounded in realistic scenarios. Its modality-aware design enables evalu- ation of both full-duplex models and modality-specific streaming models. We hope DIVE facilitates progress toward more natural and capable real-time multimodal interactive systems

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.