Certified Out-of-Distribution Safety for Generative Maps from In-Support Geometry Alone
Abstract
Every generative model is trained on a finite set of examples, and every generative model is eventually asked for something a little outside that set. What happens next is usually found out the hard way: you sample, you look, and you decide whether the output is still trustworthy. The awkward part is that checking whether a model behaves outside its training support seems to demand data from outside that support, which is exactly the data nobody has. We show the check can be done entirely from the inside. Near the edge of a model's support we fit a small atlas of local linear approximations to the map and read off three quantities that never require leaving the support: how well the atlas already fits, how rough the map looks at the finest scale the atlas can see, and how sharply it bends. Those three numbers name a radius: a distance beyond the edge within which the map's extrapolation error is guaranteed to stay under a tolerance the user picks. We test the prediction against ground truth on synthetic warped manifolds, on maps fitted to real tabular data, and on flow-matching networks trained on three image datasets. On the synthetic families the predicted radius sits inside the true one on every cell we measured. The certificate is conservative rather than optimistic, and tight enough to be useful rather than vacuous. On trained networks it does too, but only after one constant is recalibrated for the model class, and we report exactly what goes wrong when it is not. The algebra underneath is machine-checked in a proof assistant, so the parts that are theorems are theorems. Two findings surprised us. The bending term must be measured as the curvature of the map itself and not the curvature of the surface it draws. The second, more natural choice quietly smuggles in a dependence on how much the map contracts, and breaks the guarantee. And the constant tying the two roughness measurements together, genuinely constant on synthetic maps, becomes a per-model quantity with a heavy tail on trained networks, so it must be calibrated per model class and at a high quantile rather than at a typical value. We also report where the method stops: it certifies smooth deterministic samplers, and a diffusion-transformer sampler we tried falls outside that class for reasons we measure rather than guess.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.