RippleVLM: Fast Structured Visual Extraction via Constrained Block Diffusion
Abstract
Multimodal extraction needs JSON that software can parse and values that match the page. Diffusion vision language models are well suited to that contract: tokens that are already known can be written into the response and held fixed, while unknown positions are filled by denoising, with bidirectional context inside each block. For a fixed-key record the whole object skeleton is known before generation. Template-Anchored Diffusion Decoding (TADD) therefore places field names and delimiters in a response template, denoises only the bounded value spans, and serializes those values into JSON. The returned object contains the requested keys because it was assembled from the template, and a generation cap cannot cut the object off mid-structure. Diffusion-Aligned Visual Projection (DAVP) adapts only the visual projector to this template, yielding models we call RippleVLM. On multiple datasets, RippleVLM-8B returns valid JSON with every required key in 2.50 s and raises content token F1 from 40.3% to 55.0%. A autoregressive model of same size baseline with Outlines reaches 58.4% F1 but passes the same JSON check at 38%, in 4.62 s under a 256-token cap. As the schema grows from 2 to 8 fields, TADD stays at 100% JSON validity, including on DiffusionGemma-26B. These results show that fixed-template diffusion decoding produces parseable JSON for bounded extraction while leaving content accuracy as a learned problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.