SPROUT: Syntax-Progressive Denoising for Detailed Captioning in Diffusion VLMs
Abstract
Vision-language models (VLMs) are increasingly expected to produce long, detailed image descriptions, yet additional detail often comes with increased hallucination, as attributes, relations, and secondary objects are generated from linguistic priors rather than visual evidence. We attribute part of this failure to flat caption supervision, which treats sentence skeletons and fine-grained modifiers as an undifferentiated sequence. We introduce SPROUT (Syntax-PRogressive Ordered Unmasking Training), a syntax-progressive framework for detailed captioning in masked diffusion VLMs that turns the denoising axis into a within-example curriculum. We construct a dataset in which each human-written caption is decomposed into five nested levels, yielding a syntactic tier for every word. During training, a tier-conditioned corruption schedule masks fine-detail tokens earlier than skeleton tokens, shifting detail supervision toward contexts where the ground-truth sentence skeleton is visible. SPROUT further incorporates schedule-consistent weighting, same-tier block masking, and termination supervision, while requiring no architectural changes or additional inference-time computation. Across hallucination benchmarks (CHAIR, AMBER, THRONE, and OpenCHAIR) and detailed-captioning benchmarks (Image Paragraph, IIW-400, LN-COCO, and DOCCI-test), SPROUT improves fine-grained faithfulness over both the pretrained backbone and a same-data uniform-SFT baseline while preserving coverage and descriptive quality. It also retains general multimodal capability on MME and MMBench. Tier-stratified analysis further reveals emergent skeleton-first generation under the unchanged confidence decoder.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.