acceptodds
Under review as a conference paper at ICLR 2027

Stop on Sight: Omnimodal Standalone Full-Duplex Interaction Supporting Real-Time Visual Interruption

Abstract

Omnimodal full-duplex interaction is pivotal for enabling natural, human-like dialogue in interactive artificial intelligence systems. However, existing full-duplex models trigger interruptions almost exclusively via acoustic cues, halting generation only when speech barge-in is detected, while critically lacking vision-driven interruption. In practice, this capability is vital, as humans naturally cease speaking upon noticing urgent visual cues, such as an explicit stop gesture or an abrupt change in the scene. To address this fundamental limitation, we present VIFI, which, to our knowledge, is the first visual-interruptible standalone full-duplex interaction model. Operating on 400-ms streaming chunks, VIFI introduces a decoupled four-token interaction protocol that allows text emission and audio rendering to proceed at independent rates, ensuring a natural and stable speaking cadence. To support effective training under this paradigm, a robust construction pipeline is built with a dedicated focus on vision-driven interruption scenarios. Furthermore, we design a tailored reinforcement learning framework that further improves full-duplex turn-taking and interruption behavior. To facilitate comprehensive evaluation, we establish VIBench, to our knowledge the first benchmark specifically designed to assess visual interruption fidelity under continuous streaming settings. Experiments show that VIFI substantially outperforms existing full-duplex baselines, surpassing MiniCPM-o 4.5 by 40.4 percentage points in visual interruption accuracy under continuous streaming settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.