acceptodds
Under review as a conference paper at ICLR 2027

Preserving What Was Said: Evidence-Guided Audio-Visual Speech Enhancement

Abstract

Generative audio-visual speech enhancement can improve acoustic quality while increasing speech-recognition errors. Under severe noise, we observe a cleaner-but-wrong regime in which a frozen audio-visual diffusion enhancer improves acoustic quality yet increases word error rate (WER) from 0.335 for the noisy input to 0.504 after enhancement. To mitigate this failure without retraining the enhancer, we propose AVEC (audio-visual evidence consistency), an inference-time framework that evaluates multiple enhancement candidates against complementary evidence from the original observation. AVEC compares each candidate transcript with visible articulation and with the noisy-input transcript, and converts these evidence scores into input-dependent mixture coefficients for waveform aggregation. On the primary evaluation, AVEC reduces WER from 0.504 to 0.313 while keeping acoustic quality close to that of the original enhancer. The higher-budget AVEC+ further reduces WER to 0.279, below the noisy input, while improving reference-based quality metrics. A human listener study further demonstrates the effectiveness of AVEC and AVEC+.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.