How Much to Trust the Eyes? Reliability-Guided Inference for Diffusion-Based Audio-Visual Speech Enhancement
Abstract
Diffusion-based audio-visual speech enhancement (AVSE) iteratively recovers speech from noisy audio conditioned on temporally aligned visual cues. Its real-world deployment may confront distribution shifts, such as visual degradation and audio-visual mismatch, which misguide the reverse diffusion process. Our analysis discloses that visual quality alone does not necessarily determine whether visual conditioning remains useful or misguides enhancement since its effect is coupled with both the pretrained model and the acoustic input. Motivated by these observations, we propose Reliability-Guided Inference (RGI), which takes the departure of both internal visual representations and model responses from those of matched audio-video pairs as evidence of distribution shift inside the pretrained model and converts this evidence into a reliability score that regulates visual conditioning during reverse diffusion. Specifically, at each reverse step, RGI weights the difference between audio-visual and audio-only predictions according to the reliability score, without updating the pretrained models. To retain visual cues that can remain useful at low input signal-to-noise ratio, RGI imposes a lower bound on the conditioning coefficient, the conditioning floor. A diffusion-step schedule further modulates the retained visual contribution along the sampling trajectory. Experiments across four corpora and three frozen diffusion AVSE backbones show that RGI improves enhancement under a range of distribution shifts on the corpora with clean references, with gains that depend on the backbone and corpus. Recordings collected from the Ameca humanoid robot further show that RGI can be deployed for robot interaction, with demonstrations available at https://anonymous.4open.science/w/ameca-demo/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.