acceptodds
Under review as a conference paper at ICLR 2027

What Does an OOD Split Measure? Distance, Identifiability, and Completion Priors

Abstract

Out-of-distribution (OOD) benchmarks can be difficult for two different reasons: a test point may be far from the training inputs, or the training data may not contain enough information, under the chosen representation, to determine its prediction uniquely. We separate these cases on four dense six-component catalyst libraries. For each pair of elements , we hide every measured composition containing both elements and ask models to predict those held-out compositions from data in which and are observed separately but never together. A model-class-specific null-space audit shows that these test compositions can remain inside the training convex hull even though the – interaction is exactly unidentified; conversely, more distant composition-cap splits can remain fully identifiable. The resulting missing direction must therefore be filled by each estimator's inductive bias, or completion prior. An anchored difference transducer (ADT) recovers about 55–60% of the variance of the hidden interaction. Mechanism tests show that this recovery comes from a smooth separable model of property changes across composition space, rather than from low-rank structure in the pair-interaction matrix or lookup of previously observed moves. Exactly separable exponential-polynomial surfaces provide a positive control, while non-separable contamination rapidly removes the advantage. Finally, we test whether this structural match can be detected without using the true censored labels. A training-only “shadow censoring” diagnostic substantially lowers real-data model-selection regret relative to ordinary in-support holdout, although neither diagnostic reliably identifies the best method for each individual pair. The results distinguish geometric distance, design identifiability, and statistical predictability, and show that a formally unidentified interaction can still be predictable when the observed data contain transferable structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.