acceptodds
Under review as a conference paper at ICLR 2027

When Conditional Generators Cannot Extrapolate: An Exactly Measurable Decomposition of Out-of-Distribution Error

Abstract

Conditional generators answer requests they never encountered in training, but when an answer fails we cannot tell whether the model is inadequate or the target was unreachable from the available data. A recent theory separates this failure into target geometry, representation distortion, and training fit. Its sharpest criterion, however, relies on a proof step that controls a different, alignment-dependent quantity, while its experiments estimate every term through proxies. The theory therefore neither certifies the claimed boundary nor reveals which cause actually binds. We repair it by separating an unconditional certificate from a sharper criterion that holds under coherent alignment, formalizing every statement in a proof assistant, and turning distance from the training hull into a test of what any inference-time reweighting can reach. We then construct two benchmarks-capital-letter point clouds and single-cell perturbation populations-where the distances, barycenters, error terms, and sampling floor are computed exactly under held-out evaluation. In both domains, geometry predicts the unseen-condition error while representation distortion and training fit do not bind. Reweighting, condition mixing, and representation regularization leave the geometric gap in place; one genuine target example collapses it. The implication is a practical triage rule before training: determine whether a request is already reachable, requires new examples, or lies beyond what tuning alone can recover.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.