From Spatial Evidence to Target Understanding: Aligning Diffusion and Grounding Models for Reliable Visual Grounding
Abstract
Visual grounding maps an image and a referring query to a bounding box. Currently, text-to-image diffusion models appear well suited to augment this ability because of their dense visual features. However, we find that widely adopted diffusion augmentation can be harmful for visual grounding, which may perform even worse than noisy inputs. To diagnose this problem, we examine spatial evidence in model features, and find that naively generated images lead the model to disperse focus over the whole image, whereas noisy inputs coincidentally align with the grounding model’s target-centered focus, yielding higher performance. This suggests that only aligned information from diffusion models benefits visual grounding. Motivated by this finding, we propose DiReCT, a diffusion-supervised framework that strengthens vision-grounding alignment from both the diffusion and grounding sides. It employs a frozen diffusion model as a teacher with a residual-feature mechanism to extract better-aligned features, and a grounding model as a student that uses paired valid and counterfactual queries to filter out misaligned diffusion-augmented features. The student then learns from the frozen teacher to capture detailed visual information, while only the VLM student is retained at inference. Extensive experiments on GroundingME, FineCops, gRefCOCO, LISA, and RefCOCOg show that DiReCT consistently improves diffusion-augmented visual grounding and outperforms state-of-the-art methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.