From Tokens to Canvases: Rethinking Dense Layout-to-Image Generation in Remote Sensing
Abstract
Layout-to-image generation controls image synthesis through object categories and bounding boxes, yet existing instance-centric conditioning becomes difficult to scale in dense scenes, particularly in remote-sensing imagery where numerous objects often coexist. Practical capacity limits force excess object conditions to be truncated, creating a condition–image mismatch when omitted objects remain visible in the training images, which later manifests as missing or spurious objects during inference. We revisit this representation by rasterizing all bounding boxes onto a fixed-resolution global canvas, avoiding object-wise growth in the spatial conditioning input. Surprisingly, this simple representation achieves comparable or even better layout consistency than richer edge- and shape-based conditions. Building on this finding, we propose DenseGen, a category-centric layered conditioning framework that decouples category-level semantic conditioning from instance-level spatial control. Category-Level Layout Mask Attention injects shared category semantics into corresponding spatial regions, while Instance-Aware Spatial Adaptive Layering separates overlapping boxes to reduce geometric ambiguity. Experiments on DIOR-R, DOTA-v1.0, and SODA-A demonstrate improved layout consistency and semantic accuracy in dense scenes, as well as greater utility of generated images for downstream object detection, without requiring additional structural annotations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.