acceptodds
Under review as a conference paper at ICLR 2027

RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning

Abstract

Dense image captioning is critical for vision-language pretraining and for text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthesizing captions from strong vision-language models (VLMs) is a practical alternative, supervised distillation often yields limited output diversity and weak generalization. Reinforcement learning (RL) could overcome these, but its successes have so far been concentrated in verifiable domains that rely on deterministic checkers—a luxury not available in open-ended captioning. We address this verification bottleneck with RubiCap, an RL framework that derives fine-grained, sample-specific rewards from LLM-written rubrics. For each image, an LLM rubric writer compares captions from a diverse committee of VLMs to identify consensus strengths and diagnose the current policy's deficiencies. These findings are converted into explicit evaluation criteria, enabling an LLM judge to decompose quality assessment and replace coarse scalar rewards with structured, multi-faceted assessments. Across extensive benchmarks, RubiCap achieves the highest win rates on CapArena, outperforming supervised distillation, prior RL methods, human-expert annotations, and GPT-4V-augmented outputs. On CaptionQA, it shows superior word efficiency: our 7B model matches Qwen2.5-VL-32B-Instruct, and our 3B model surpasses its 7B counterpart. Remarkably, using a compact RubiCap-3B as a captioner produces stronger pretrained VLMs than those trained on captions from proprietary models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.