Prompt-Switchable Topological Vectorization: Structured Geometry Decoding from a Frozen Segmentation Foundation Model
Abstract
Downstream geospatial applications often require segmentation outputs to be represented as polygons satisfying hard geometric and topological constraints: one polygon per instance, no overlapping interiors, and exact shared boundaries between adjacent polygons. Converting masks to polygons is usually left to post-processing, which makes topology something to repair rather than something to guarantee. We move these constraints into the decoder: a frozen promptable segmentation backbone is adapted with LoRA and a geometry head on the same features, M trainable parameters versus M frozen, and the predicted geometry fields are decoded into a planar graph whose faces satisfy these constraints by construction. Our experiments reveal two broader findings. Instance metrics are close to blind to structural decodability: dropping the geometric supervision costs half a point of micro F1 while raising uncovered instances from 11.3% to 96.7%. And an open vocabulary at the prompt does not reach the output: switching the prompt alone gives micro F1, whereas ten annotated images give . On three cropland domains, the topological output outperforms all evaluated baselines, with a -point gain on the one genuine zero-shot transfer, while returning ten to sixteen times the shared boundary of a contour pipeline, and still seven to nine times after that pipeline's output has been through a standard GIS repair.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.