Q-VG: Towards Quantitative Reasoning for Weakly Supervised Visual Grounding
Abstract
Weakly supervised visual grounding (WSVG) aims to locate objects referred by natural language queries, without region-level human annotations. However, existing methods struggle to ensure the correct number of targets, a common requirement in compositional reasoning. In this paper, we propose an end-to-end network, Q-VG, for quantitative reasoning and grounding under weak supervision. Specifically, Q-VG incorporates a novel Referring-Object Counting (RECO) module to estimate the number of objects referred to by the input expression. It enhances quantitative awareness prior to grounding and determines the number of bounding boxes to regress. Besides, we also propose a novel weakly supervised objective, which encourages multiple positive objects to align with the given referring phrase. To further evaluate the ability of quantitative reasoning, we introduce QGround-Bench, a quantitative grounding benchmark. It comprises three subsets, Exact-QG, Gen-QG and Comp-QG, which progressively evaluate exact quantitative constraints, generalized quantitative expressions, and quantitative compositional reasoning. Experiments show that Q-VG achieves competitive performance on both the proposed and classical benchmarks. The source code and dataset will be available on GitHub after double-blind phase.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.