Situated Inference: Adapting Frozen Generative Speech Enhancement under Acoustic Distribution Shift
Abstract
Diffusion and flow matching models typically enhance speech using fixed inference strategies that may become unsuitable under acoustic distribution shifts in real-world deployment. Our cross-condition analysis shows that the advantage of an inference strategy is not consistently preserved across acoustic conditions, even when the model weights and solver update rule remain fixed. This finding motivates Situated Inference, which specializes the inference strategy of a frozen generative enhancer to its situation, the target acoustic distribution under which it is deployed. Using paired target speech, Situated Inference evaluates a structured inference space that jointly varies observation processing, the reverse starting point, and the solver step count, thereby constructing a target-specific quality and computation profile. From this profile, it selects one strategy either to maximize enhancement quality or to minimize computation subject to the default quality, then fixes that strategy for subsequent inputs. Across four generative backbones and three acoustic conditions, Situated Inference improves perceptual speech quality under quality specialization and reduces the number of function evaluations under efficiency specialization. Deployment on the Ameca humanoid further illustrates the efficiency advantage of Situated Inference in practice, with demonstrations available at https://anonymous.4open.science/w/situated-inference-demo-0761/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.