acceptodds
Under review as a conference paper at ICLR 2027

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

Abstract

Visual captioning converts visual information into text, requiring both completeness of salient content and correctness of expressed claims. Multimodal large language models (MLLMs) have advanced captioning through scaling and highquality data, while reinforcement learning (RL) offers a way to refine these factual dimensions. However, directly judging captions with an MLLM can yield coarse quality scores, while visual question answering (VQA) rewards supervise only queried facts. Both routes struggle to provide fine-grained rewards for individual omissions and errors. We propose VCap, a Witness–Adjudicator factcomparison reward that uses reference facts as inspection cues and visual evidence to adjudicate omissions and contradictions. References may be incomplete or imperfect; rewarding supported details beyond them enables weak-to-strong learning. A unified hypergeometric model explains how partial factual checks detect both coverage deficits and erroneous assertions under explicit sampling assumptions. Global and local temporal comparisons extend the same reward protocol from images to videos. Training Qwen3-VL-8B with VCap yields the highest aggregate scores among the evaluated models on CapMAS, DecapBench, and VDC. Ablations support the complementary reward components, and human assessment supports more complete captions. RL also outperforms backbone best-of-8 distillation, while the learned captions and policies benefit supervised training and visual question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.