CuRe: Claim-Level Rubric Rewards for Video Captioning
Abstract
Existing reward functions for dense video captioning largely assume reference captions provide complete supervision, overlooking that a single reference can rarely cover all valid visual content in a video. To address this issue, we propose Claim-Level Rubric Rewards (CuRe), a reward framework that decomposes captions into visual claims such as entities and actions, and verifies each claim directly against the video. This claim-level verification enables CuRe to distinguish fine-grained errors from valid details omitted by the reference. We then define a reference-anchored credit weighting scheme that assigns credit based on both visual support and reference alignment: reference-matched claims receive full credit, video-supported claims beyond the reference receive bounded credit, and unsupported claims receive zero credit. This design rewards visually grounded coverage rather than strict reference matching in text, allowing valid details beyond the reference to contribute while keeping unsupported claims penalized. On three public captioning benchmarks, CuRe achieves the strongest results among open-source models. To test whether CuRe-generated captions provide better supervision than the original captions, we re-caption three video-caption corpora (867K samples) and pretrain VLMs with either the original or CuRe-generated captions. CuRe-recaptioned data improves performance on 25 of 27 metrics across six downstream video-understanding benchmarks, demonstrating that CuRe-generated captions provide stronger video-language supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.