DECAR: Evidence-Grounded Clustered Calibration for Fine-Grained Vision–Language Alignment
Abstract
Fine-grained image–text scoring becomes difficult when captions differ only in a local attribute, relation, count, or factual detail. A model may rank such captions correctly while still assigning scores that disagree with the reference scale. We introduce DECAR, a continuous vision–language scoring framework trained on image-centered groups of graded captions together with element-level support. DECAR combines a distributional global score head, element-level prediction, within-image ranking constraints, and soft evidence bounds that act on implausible global scores during training. The element branch is removed at inference, leaving only the image–caption branch and global score head. On the Qwen2.5-VL-7B instantiation, DECAR achieves 0.7583 SRCC and 0.5864 MAE. In a separate prompt-disjoint frozen-CLIP study with matched training settings, clustered labels increase mean SRCC from 0.4660 to 0.5601 across three seeds. Target-shuffle and gap-aware controls further show that caption-specific graded labels mainly improve correlation, whereas additional ranking constraints reduce numerical error. Additional analyses examine paraphrase stability, counterfactual ordering, residual errors, and performance across semantic subsets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.