Release-State Audits for Retrospective Machine Learning
Abstract
A dataset name or event-time cutoff does not uniquely specify the evidence available to a retrospective machine-learning evaluation. We introduce release-state audits: a closed reference specification linking exact source-state commitments, support accounting, estimands, and admissible claims. It separates support migration from retained-unit change and specifies typed non-results when identity, label, or missing-support requirements fail. The current v0.2 one-input checker validates only closed aggregate types and listed algebra over caller-supplied declarations and always withholds; it does not establish source replay or authorization. Three formal boundaries motivate the specification: a terminal event-time view need not identify earlier availability; finite agreeing label states do not establish permanence; and union-paired contrasts with release-exclusive support depend on registered missing-side restrictions. A separate closed synthetic adapter/evaluator fixture supports only finite implementation conformance, not real-world detection performance. We examine one exploratory pre-merge BIG-bench task-state case using three exact model revisions and a fixed census of 998 effective inputs. The complete four-cell panel exhibits descriptive grade differences between task states. Under no-appended-choice rendering, all three models have identical aggregate grades near the cell-specific uniform-choice expectation, without establishing identical predictions. This limited vignette illustrates why task state belongs to the evaluation object; it does not establish post-release benchmark drift, causal repair effects, or general model capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.