Spatial Residual Adaptation of Frozen Audio-Visual Speaker Extractors
Abstract
Single-channel audio-visual target-speaker extractors leave the spatial information of microphone arrays unused. We propose D-CRPF, a 70.9K-parameter dual-direction complex residual post-filter that adds spatial evidence to a frozen extractor without accessing its hidden features or training graph. Two analytic beamformers condition complex reference correction and signed beam-disagreement correction around the frozen output. We characterize identity initialization, input-dependent correction bounds, and reference-preserving channel-permutation invariance. On simulated six-microphone mixtures covering all 1,243 full-duration official LRS2 test utterances, D-CRPF improves matched AV-CrossNet by 1.254 dB SI-SDR versus 0.045 dB from an equal-parameter single-channel control, reaching 14.433 dB and 8.14% WER. Under a fixed-reference circle-to-line change, the complete system loses 0.806 dB versus over 4 dB for direct multichannel references; its gain over Stage 1 remains positive but contracts from 1.246 to 0.441 dB. The residual-only intermediate variant, CRPF-R, isolates the first correction direction and provides complementary transfer evidence: it improves three released front ends by 0.904–1.502 dB SI-SDR and matched IIANet by 1.561 dB. Missing-video tests expose failure boundaries.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.