acceptodds
Under review as a conference paper at ICLR 2027

DQ-RL: Fine-Grained Rewards from Diverse Queries for Dense Image Captioning

Abstract

Image captioning is a fundamental task in vision-language understanding and a core capability of large vision-language models. Existing efforts to improve this capability mainly rely on supervised fine-tuning of base models, often on domain-specific data, while recent work has begun to explore reinforcement learning (RL) for better generalization. However, applying RL to image captioning remains challenging: as an open-ended generation task, captioning lacks verifiable rewards, making reward design under the Reinforcement Learning with Verifiable Rewards (RLVR) paradigm particularly difficult. To address this challenge, we propose **Diverse-Query Reinforcement Learning** (**DQ-RL**). We design completeness queries (**CQs**) and faithfulness queries (**FQs**), perform unified sampling over the candidate query pool, evaluate sub-capabilities of generated captions, and use the resulting quantitative scores as rewards for optimization. Following this paradigm, we construct a k training dataset named **DQ-CAP-40k**. Extensive experiments show that DQ-RL, a 4B model, achieves an average improvement of \(1.4%\) over baseline models under the PRISM evaluation framework, and matches Qwen3.5-27B model at \(63.1%\). We also build **DQ-Bench**, an evaluation benchmark tailored for image captioning. On the three core metrics of DQ-Bench, namely Correct, Omit and Wrong, DQ-RL outperforms the baseline by , , and , respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.