acceptodds
Under review as a conference paper at ICLR 2027

Auditing Audio–Visual Integration in Omni-Modal Models

Abstract

Video can change speech recognition without showing whether a model uses the visible articulation. We evaluate eight omni-modal models using silent word recognition, speech in noise, repeated and permuted frames, and cross-dubs that hold the auditory waveform fixed across visual donors. The six open-weight models recognise silent command words at near-chance accuracy and show little average movement toward the donor's word, although their output distributions vary across donors. The two closed-weight models benefit from native video in selected noisy conditions. A follow-up tests one of these models with matched waveforms across native, reprocessed and dubbed clips and finds no clear benefit or preference for compatible video. A masker-only tail changes its audio-only accuracy by about eleven percentage points. A specialist audiovisual recogniser shows donor-directed effects on the same files; its benefit survives reprocessing and depends on the donor's word. We report word and onset outcomes separately, including transcripts that the onset rule cannot classify. Exploratory probes recover label information after linear nuisance adjustment, but the tested interventions do not establish how that information affects recognition. These results distinguish changes caused by adding video from evidence that a model uses visual speech, while leaving their architectural and training causes unresolved.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.