Grounding as Verification: Closing the Loop in Grounded Video Captioning
Abstract
Grounded video captioning systems generate descriptions with phrase-level localizations in the video. However, existing pipelines still flow one way, caption grounding: grounding reveals where the caption points, but never feeds back whether generated entities are visually supported. We push this view further by making grounding act as a verifier, forming a caption grounding loop in which localization evidence verifies and refines caption generation. For this purpose, we build the closed loop from two complementary directions. First, we observe that current grounding heads typically return only a single representative trajectory per phrase, regardless of whether the phrase refers to one or multiple entities. This provides only weak, presence-level verification: a plural phrase should be supported by multiple visual instances, yet the grounding head exposes only one. We therefore introduce a lightweight Multi-Instance Adapter that expands each phrase query into multiple persistent evidence slots, upgrading grounding into a structured multi-instance verifier. Second, we build annotation-free visual self-alignment with instance-presence and temporal-consistency rewards for KL-regularized policy optimization, improving captioning and grounding without ground-truth captions or extra box supervision during self-alignment. On iGround, the adapter improves AP50 from 40.0% to 51.7% over GROVE, and the full closed-loop model reaches 53.3% AP50 and 88.0 CIDEr. Results on VidSTG and ActivityNet-Entities further show effectiveness beyond the main benchmark. Together, these results support a verification-centered loop in which grounding provides feedback for better captions, and improved captions induce cleaner grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.