acceptodds
Under review as a conference paper at ICLR 2027

Identifiability Precedes Generalization: A Pre-Fit Audit for Target–Domain Aliasing

Abstract

When the target label is determined by the acquisition domain, no amount of internal predictive performance establishes target-specific attribution from the observed data alone. Every cross-validated score computed from the labels equals the one computed from the domain-induced labels, so the competing explanations for the same number are indistinguishable. Distribution-shift evaluation asks how well a model trained on one distribution performs on another. This audit asks a different and prior question: what the recorded design allows a result to be attributed to. The two are complementary. The audit has a pre-fit structural part and a model-based diagnostic part. Structurally, under an additive target–domain model the classical connectivity of the bipartite support graph identifies which target main-effect contrasts are estimable, including in designs that are incompletely crossed, where some class never appears in some domain. The singular values of the residualized target basis measure proximity to aliasing, which the rank alone cannot. Diagnostically, a classifier two-sample test conditioned within a target stratum tests for domain-associated differences the target model could exploit. We report its operating characteristics and the shifts it misses. For a binary target and a score carrying no target information beyond the domain, we give an exact identity, valid at any class prevalence, between that score's AUC against the target and its AUC against the domain, given the class–domain allocation. On three datasets the audit separates different threats. On a sensor benchmark every additive target main-effect contrast is estimable despite incomplete crossing, yet random splitting reports accuracy 11.2 percentage points higher than leave-one-batch-out evaluation. In our Camelyon17 analysis the design is sound and hospital separability is nonetheless high: strong domain information and estimable target contrasts coexist. Grouping folds by slide lowers the reported AUC more than the subsequent hospital holdout does. In a clinical breath cohort disease status is inseparable from acquisition series, and two healthy series separate at AUC=1.00. The audit tells a study which contrasts its design supports and what its evaluation must block on. Applied to our own analysis it withdrew a date-prediction result, under random folds and 0.48 once folds are grouped by collection date.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.