acceptodds
Under review as a conference paper at ICLR 2027

SatLen: Evaluating Semantic-to-Satellite Translation Across Nested Scales

Abstract

Image generators are increasingly asked to render the world from a map. We ask whether they respect a map's physical scale. We introduce SatLen, a benchmark that renders the same 100 scenes from 45 countries and territories as land-cover maps at four ground footprints, from to  km per side, at a fixed input resolution. Across 21 generator configurations and 8,400 generations, we find a coverage paradox: as the footprint widens, every configuration becomes more similar to the real image under CLIP while matching the requested layout less closely. After removing the overlap that correct class proportions earn by chance, spatial placement falls by , and three configurations fall to near-chance placement at  km. Pixel-level diagnostics that bypass the segmenter show where the failure lies. Generated edges still follow the map's boundaries, and do so more sharply at wide coverage, yet the land-cover identity of the enclosed regions is lost. A segmenter-free nearest-class-mean classifier confirms that this semantic loss is not a segmenter artifact. Models also treat scale as a texture setting: output textures are too smooth close up and gain fine detail far faster than real imagery as coverage grows. A prompt ablation without the extent sentence shows the same ramp, indicating that models read scale mainly from the map. Finally, a model's close-up and wide views of the same location rarely depict the same place. Current generators keep the lines but lose the land.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.