EAG-Cap: Evidence-Aware Grounded Caption Revision with Visual Evidence Screening and Caption-Driven Grounding Recovery
Abstract
Grounded vision-language models connect generated descriptions to image regions, but dense grounded proposals mix useful localized content with unsupported or overly specific claims. Conservative visual evidence screening can clean these proposals, yet it introduces a precision–coverage trade-off: valid caption-relevant regions may be pruned together with noisy ones. We present EAG-Cap, a training-free framework that treats proposal phrases as candidate evidence-bearing claims rather than trusted facts. EAG-Cap screens their localized visual support, composes a concise caption under structured evidence constraints, and then uses that final caption to identify groundable concepts that still lack compatible spatial support. Caption-driven grounding recovery selectively re-queries only these missing concepts and merges recovered regions with retained grounding. On the full 5,000-image COCO val2017 split, EAG-Cap obtains 0.839 CIDEr and 18.01% CHAIRs, reducing hallucination relative to Florence-2 Detailed (29.62% CHAIRs) while retaining more annotated content than similarly concise baselines. More importantly, recovery nearly restores the caption-conditioned coverage of raw proposals (90.97% vs. 90.96%) while retaining 84.54% [email protected], compared with 61.45% for raw proposals, and reducing boxes per image from 19.75 to 4.91. A controlled text-only rewrite explains part of the language behavior but provides no caption-aligned grounding output. These results characterize EAG-Cap as a coupled evidence-revision and grounding-recovery framework rather than a caption filter alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.