Beyond Visual Bounds: Generation-Aware Layout Analysis for Editable Visual Documents
Abstract
Vision-language models (VLMs) have advanced text recognition and document parsing. However, when a parsed text element is edited, bounding boxes that capture only its visible extent do not specify the space available for replacement text or how that text should be aligned within that space. We introduce Generation-aware Layout Analysis (GLA), a task that recovers text, allocated regions, semantic categories, and alignment to provide spatial constraints for subsequent element-level visual editing. This representation could support editable slide and document reconstruction, translation with layout preservation, and automated design editing. For training and evaluation on this task, we construct a 37K-image training dataset spanning nine domains and introduce GLA-Bench, a 1,000-image benchmark with four tasks. To predict dense coordinate and attribute sequences accurately and efficiently, we introduce joint supervised fine-tuning (SFT), reinforcement learning (RL) with layout-specific rewards, and accelerated decoding. To recover local feedback obscured by image-level aggregation, we introduce EAA-GRPO, which augments group-relative policy optimization (GRPO) with Element-level Advantage Assignment (EAA). EAA compares predictions for the same element across rollouts to provide field-specific supervision. To accelerate decoding, we adapt a block-diffusion drafter to the repetitive structure of parsing outputs. On GLA-Bench, our 4B parser reaches 46.59 [email protected] in full parsing, versus 6.26 for the best evaluated zero-shot baseline. Layout-only RL improves full parsing on both backbones. Task-adapted drafting achieves 2.18 autoregressive throughput with the SFT target fixed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.