Language-Conditioned Mean-Field Unfolding: A Variational Inference Paradigm for Referring Expression Segmentation
Abstract
Deep algorithm unfolding bridges iterative optimization and differentiable neural architectures, yet its formulation has largely centered on recovering latent coefficients under predefined structural priors. In this work, we extend this paradigm to language-conditioned semantic inference, where the object to be unfolded is not an image-reconstruction coefficient, but a semantic posterior over a query-conditioned probability simplex. We propose Language-Conditioned Unfolding of Mean-Field Inference Networks (LUMIN), which formulates referring expression segmentation (RES) as variational mean-field inference in a text-conditioned energy model. LUMIN dynamically constructs a semantic state space from textual hypotheses together with an adaptive background state. A text-derived Gram compatibility kernel is coupled with a symmetric visual affinity graph to propagate language-conditioned evidence through simplex-constrained Softmax updates. Unrolling this inference process yields a compact and interpretable network in which each stage corresponds to a principled surrogate step on the variational free energy. A background-relative differential readout further calibrates marginal log-odds by suppressing encoder-induced common-mode drift. Experiments on RefCOCO/+/g demonstrate competitive segmentation performance and transparent step-wise posterior refinement with only three unfolding stages.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.