acceptodds
Under review as a conference paper at ICLR 2027

Scientific Image and Mask Generation with Adaptive Patching

Abstract

Building annotated datasets of high-resolution scientific images faces two bottlenecks: computationally expensive image synthesis and labor-intensive pixel-wise annotation. We present a unified framework that addresses both through a shared, spatially adaptive latent representation for image–mask synthesis. An edge-guided quadtree allocates finer patches to structurally complex regions, and a shared codec maps variable-size patches into a common latent space. Conditioned on reference layouts, a rectified flow model generates latent features from which image and semantic mask branches decode paired outputs. The mask branch is trained using simulator-provided labels with the image codec frozen, avoiding manual annotation for its training supervision. On concrete X-ray CT data, our framework generates images at resolutions up to 8192 8192, achieving a global FID of 18.04 and a native-resolution patch-FID of 22.50. At 4096 4096, it achieves global and patch FID scores of 17.98 and 40.98, compared with 46.84 and 73.73 for the evaluated class-conditioned dense rectified-flow baseline, alongside approximately 2.3 higher generative-model training throughput. Evaluation against manual annotations on 12 generated images yields a three-class macro Dice of 0.736. These results establish the feasibility of combining efficient high-resolution generation and semantic decoding within a single framework for scientific image–mask synthesis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.