TidyUp: Post-Alignment of Agentic Layout Outputs with a Tiny Specialist
Abstract
Large language and vision–language agents can now assemble and edit real graphic designs, but they reason about geometry imprecisely: leaving overlaps, broken alignments, uneven spacing, and off-canvas content. The prevailing fix—re-invoking the same or a larger model to critique a rendered canvas and rewrite coordinates—is slow, costly, non-deterministic, and prone to regressing already-correct parts of the design. We argue that this narrow, geometric defect class does not call for more generalist reasoning but for a specialist. We present TidyUp, a tiny (86M-parameter), vision-free pipeline that repairs a design in place after the agent has already done the visual reasoning. A lightweight reflection router first flags which elements are broken and by what violation. Specialist correctors — first to jointly condition intricate design group structures and prior design state — apply hierarchical attention over local and global alignment geometry to nudge flagged elements into place, leaving size, type, and unedited elements untouched, with no regeneration. Unlike prior flow-based layout models, trained on noise-to-data, each corrector is trained for refinement, integrating its ODE from the broken layout. Owing to varying repair severity, TidyUp iterates reflection-correction with an auto-stop. Correctness-preserving guards—including a violation-regression bailout—guaranties TidyUp never returns a layout worse than its input. TidyUp, trained on AI-edited designs from a large-scale production system and a public Crello benchmark, with synthetic defects derived from multi-turn editing patterns, matches or exceeds the SOTA-level performance at a fraction of the cost: 1.2s latency at 250 RPM on a single H100, 160 MB footprint, with no additional agent calls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.