acceptodds
Under review as a conference paper at ICLR 2027

Dynamics-Guided Distribution Shaping for Reliable Vision-Language Grounding

Abstract

Vision-language grounding localizes visual regions described by natural language, yet current systems rarely indicate when their localization predictions are unreliable. Existing single-pass probabilistic regression methods mainly estimate uncertainty from the current prediction state and therefore overlook how sample-wise grounding responses evolve during optimization. We propose Dynamics-guided Distribution Shaping (DDS), which uses inter-epoch prediction variation between two consecutive epochs as a local training-dynamics signal. DDS converts this variation into a sample-specific reliability coefficient that guides bounded prediction-distribution shaping and strengthens residual correction for unstable samples. The dynamics branch is used only during training, preserving single-forward-pass localization and uncertainty estimation at inference. Across RefCOCOg and RefCOCO+ with three multimodal backbones, DDS consistently improves uncertainty–error alignment while maintaining competitive grounding performance. Under zero-adaptation transfer from RefCOCO+ to Ref-Adv-s, DDS further improves uncertainty reliability and out-of-distribution discrimination under semantic and reasoning shift.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.