acceptodds
Under review as a conference paper at ICLR 2027

RippleVLM: Fast Structured Visual Extraction via Constrained Block Diffusion

Abstract

Multimodal extraction needs JSON that software can parse and values that match the page. Diffusion vision language models are well suited to that contract: tokens that are already known can be written into the response and held fixed, while unknown positions are filled by denoising, with bidirectional context inside each block. For a fixed-key record the whole object skeleton is known before generation. Template-Anchored Diffusion Decoding (TADD) therefore places field names and delimiters in a response template, denoises only the bounded value spans, and serializes those values into JSON. The returned object contains the requested keys because it was assembled from the template, and a generation cap cannot cut the object off mid-structure. Diffusion-Aligned Visual Projection (DAVP) adapts only the visual projector to this template, yielding models we call RippleVLM. On multiple datasets, RippleVLM-8B returns valid JSON with every required key in 2.50 s and raises content token F1 from 40.3% to 55.0%. A autoregressive model of same size baseline with Outlines reaches 58.4% F1 but passes the same JSON check at 38%, in 4.62 s under a 256-token cap. As the schema grows from 2 to 8 fields, TADD stays at 100% JSON validity, including on DiffusionGemma-26B. These results show that fixed-template diffusion decoding produces parseable JSON for bounded extraction while leaving content accuracy as a learned problem.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.