OmniDuet-Bench: Do Omni-Modal Models Act on What They See in Full-Duplex Audiovisual Conversation?
Abstract
Real-time omni-modal models can now generate speech while receiving streaming audio and video, yet no benchmark systematically evaluates whether they respond to the speech, facial expressions, and actions in their audiovisual input with appropriate timing and content: full-duplex benchmarks evaluate only spoken interaction, and video understanding benchmarks assess comprehension of video content rather than how a model, as a conversational participant, responds to the user on screen. We present OmniDuet-Bench, which evaluates full-duplex audiovisual conversation along three dimensions: Interaction Capacity, Visual Perception, and Cross-Modal Integration. The benchmark comprises 707 human-annotated task instances curated from natural two-person video calls and organized into six streaming tasks. Visual Perception and Cross-Modal Integration tasks share the same expression and action events and are complemented by matched audiovisual and audio-only conditions; together, these designs dissociate the perception of a visual cue from its use in conversational behavior. We evaluate seven real-time omni-modal models and find that no model leads on all three dimensions: models often fail to hold the floor and to take the turn promptly when the user yields it, and even after identifying an expression or action correctly, they often fail to adapt their responses to it. Data and evaluation code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.