Escaping the Barycentric Ceiling: A Training-Free Diagnostic for Out-of-Distribution Conditional Generative Models
Abstract
A conditional generative model learns to produce data at the settings it was trained on, and is then asked for a setting it has never seen. A recent line of analysis gives a discouraging answer: if the model can only blend the conditions it was trained on, there is a hard floor on how well it can ever do, set by how far the unseen target lies outside the region those conditions span. We revisit that floor and find that it rests on a hidden assumption - that the blending weights must be positive. Allow them to go negative, so that the model extrapolates along the geometry of covariance matrices rather than averaging within it, and the floor stops applying. We prove this, and we correct an error we found in our own earlier argument for it along the way; the corrected proof is short enough to be checked by machine, and we check it. But escaping the floor is not always possible, and whether it is depends on how the data's shape drifts as the condition changes. If the drift is a stretching of the same axes, extrapolation works; if it is a rotation of those axes, nothing we tried helps. We turn this into a diagnostic that costs two covariance estimates and no training at all, and it predicts the answer before a single gradient step. It holds up on natural image and biological data, and we report plainly where it does not: it reads rotation-like drift correctly everywhere we looked, it is confounded on cluttered colour scenes, and it says nothing useful about unordered categories, which we document rather than hide. Finally we ask whether any of this survives an actual trained network, and here the most useful result is a negative one. Handing the network the extrapolated statistics directly - the design the theory naively suggests - destroys the diagnostic and makes the model worse than a plain baseline. Building the extrapolation into the architecture as a closed-form starting point, and asking the network only for a correction, restores it. The lesson is that the escape route has to be part of the model's structure; a network will not find it on its own.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.