Layout Beyond Format: Internalizing Spatial Execution in LLMs via Element-Level Geometric Rewards
Abstract
Large language models generate visual code—SVG, HTML, slide scripts—that is almost always syntactically valid, yet the rendered layouts frequently exhibit spatial defects; existing systems compensate through repeated inference-time detection and rewriting, at a cost that accumulates with every call. We observe that models do not lack spatial judgment: they reliably identify and follow spatial constraints, yet fail to maintain them during autoregressive generation. This local planning failure lies in the decoding process itself rather than in understanding, long context, or format syntax, and since the cause is format-agnostic, layout ability should be acquired once, at the format-agnostic level. We therefore build a two-stage framework that first generates a format-agnostic layout skeleton and then realizes it into code of any target format, making planning an independently supervisable object, and post-train on top of it: SFT distills two-stage trajectories, and RL uses a verifiable dense reward from a geometric detector that attributes every defect to the tokens that produced it, alongside a page-level outcome reward. A 9B model trained this way raises its end-to-end delivery rate from 25.6% to 85.6%; the two-stage formulation on its own, without training, does not help—it lowers the same model to 17.1%—so the gain comes from internalizing the geometric signal into the weights rather than from the scaffold. As the diagnosis predicts, the resulting layout ability transfers zero-shot to code formats absent from all training data and improves quality and efficiency inside an established agentic pipeline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.