Structural Extrapolative Data Generation: Hardness, Feasibility, and Formal Conditions
Abstract
Achieving reliable extrapolation, specifically, extrapolative data generation beyond the support of the training distribution, is a fundamental challenge and inherently requires strong assumptions. This paper aims to address this problem and proposes a framework for Structural Extrapolative Data GEneration (SEDGE) based on suitable assumptions on the underlying data-generating process. We provide conditions under which data satisfying novel specifications can be generated reliably, together with the approximate identifiability of the distribution of such data under certain “conservative” assumptions, as well as the inherent non-identifiability of this distribution without such assumptions. On the algorithmic side, we develop two methods for extrapolative data generation: structure-informed likelihood-weighted resampling (LWR) and likelihood-guided diffusion (LGD). We verify the extrapolation performance on synthetic data and also consider extrapolative property-guided molecule design and extrapolative text-guided image generation as real-world scenarios to illustrate the validity of the proposed framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.