acceptodds
Under review as a conference paper at ICLR 2027

SPROUT: Syntax-Progressive Denoising for Detailed Captioning in Diffusion VLMs

Abstract

Vision-language models (VLMs) are increasingly expected to produce long, detailed image descriptions, yet additional detail often comes with increased hallucination, as attributes, relations, and secondary objects are generated from linguistic priors rather than visual evidence. We attribute part of this failure to flat caption supervision, which treats sentence skeletons and fine-grained modifiers as an undifferentiated sequence. We introduce SPROUT (Syntax-PRogressive Ordered Unmasking Training), a syntax-progressive framework for detailed captioning in masked diffusion VLMs that turns the denoising axis into a within-example curriculum. We construct a dataset in which each human-written caption is decomposed into five nested levels, yielding a syntactic tier for every word. During training, a tier-conditioned corruption schedule masks fine-detail tokens earlier than skeleton tokens, shifting detail supervision toward contexts where the ground-truth sentence skeleton is visible. SPROUT further incorporates schedule-consistent weighting, same-tier block masking, and termination supervision, while requiring no architectural changes or additional inference-time computation. Across hallucination benchmarks (CHAIR, AMBER, THRONE, and OpenCHAIR) and detailed-captioning benchmarks (Image Paragraph, IIW-400, LN-COCO, and DOCCI-test), SPROUT improves fine-grained faithfulness over both the pretrained backbone and a same-data uniform-SFT baseline while preserving coverage and descriptive quality. It also retains general multimodal capability on MME and MMBench. Tier-stratified analysis further reveals emergent skeleton-first generation under the unchanged confidence decoder.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.