From Symbolic Mapping Failure to Numerosity Specialization in Vision-Language Models
Abstract
Visual number sense supports counting objects, comparing group sizes, and determining the size of selected or combined groups. Vision-language models exhibit sensitivity to quantity, but their ability to answer numerical questions remains uneven across visual scenes and tasks. We introduce a two-phase post-training framework to develop more accurate and reusable visual number sense in pretrained models. Configuration Coverage first associates shared count labels with complementary object layouts and appearances. A joint update then combines Relational Numerosity Supervision, which links counts through local changes and part–whole relations, with multi-interface supervision for comparison and query-conditioned counting. We also construct a controlled synthetic dataset for training and evaluation across visual configurations and count ranges. Experiments on two VLM families show substantial improvements in synthetic counting and task-dependent gains on external numerical benchmarks. Matched-budget ablations establish the benefit of complementary configuration coverage, while supervision-allocation studies demonstrate the contributions of relational and task-specific training. Representation analyses reveal more accurate intermediate-layer count readout and stronger cross-configuration readout transfer. We characterize the resulting profile as numerosity specialization: high-precision counting develops within targeted visual and numerical regimes, with varying transfer to larger cardinalities, other visual distributions, and numerical tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.