LayoutLattice: Pairwise Attention Biases for Contextual Layout Generation
Abstract
We present a simple, yet powerful mechanism to inject pairwise context into layout generation models via attention biases. We target relation-conditioned layout generation (the structure-preserving layout-variation setting), which produces alternative arrangements of page elements while preserving given pairwise relations such as which elements overlap and their reading order. Recent approaches based on transformers and continuous diffusion have advanced the field, yet they struggle to capture the pairwise relationships between page elements. Rather than relying solely on per-element features, our approach augments attention with an N x N bias matrix that injects pairwise priors (overlap/IoU and, optionally, reading order) computed from a reference layout and supplied as a conditioning signal, where N is the number of page elements. The mechanism is drop-in: binary masks and continuous priors are stacked as channels and applied as additive attention biases to a subset of attention heads, while the remaining heads stay unmasked. Across two regimes (max 25 and max 50 page items), our approach improves both realism and structure, e.g., reducing FID from 11.6 to 2.7 in the 50-item regime while increasing MaxIoU. Crucially, supplying the same priors as plain per-token inputs instead degrades the model, showing that the gains stem from the interface rather than mere access to the priors. We further demonstrate generating variations in different canvas sizes by using a relative-area (RA) representation of canvas and page items. Our results suggest that the N x N attention map is the natural representation for pairwise relational structure: a robust, reusable, parameter-free interface for structure-aware layout generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.