When Does a Predictable Residual Become a Scientific Correction? A Residual-Transport Audit for AI-Assisted Scientific Discovery
Abstract
Automated scientific-discovery systems can make the residuals of established models predictable, but predictive gain alone does not identify what has been found: a missing scientific contribution, a regime-specific offset, or structured noise. We formulate this downstream problem as correction acceptance and introduce a residual-transport audit that evaluates frozen corrections across scientifically meaningful boundaries. The central observation is that held-out-regime performance mixes two forms of evidence: improvement relative to information available before the regime is observed, and explanation of variation within that regime after its own residual level is accounted for. We derive the resulting decomposition and build it into a leakage-safe audit with frozen candidate selection, regime-level replication, uncertainty, and a measurement-noise stopping rule. Across thermodynamics, nuclear physics, mammalian metabolism, cosmology, and a supplementary asset-pricing control, the audit separates corrections that transport from patterns that merely locate regimes. A known equation-of-state correction transports across unseen fluids (R^2_pool=0.993), and a simple correction to Geiger-Nuttall predicts nuclides discovered decades later (R^2=0.635). In contrast, mammalian temperature is repeatedly selected yet fails within unseen taxonomic orders, while COBE/FIRAS requires no further correction. The contribution is an evaluation layer for AI-assisted science: candidate generation may be flexible, but scientific acceptance should be indexed by the environments in which the corrected model is expected to hold.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.