Contrast-First Calibration: Reusing Labels to Compare Synthetic Sources
Abstract
A cheap synthetic source is useful when its errors can be corrected with affordable real labels. Choosing that correction also uses labels, which should remain available for training. We study how to evaluate these choices when the labels are reused. Our central protocol conditions on a finite, paid source bank and allows its data to determine the correction subspaces. For hard corrections with a common real-data count, a pairwise risk error separates into two random quantities and a deterministic retention factor. One confidence event therefore covers every later real-data budget with the same radius as the corresponding fixed-budget construction. The assessment extends to unknown Gaussian noise variance through regression residuals. For a broader pooled-data selector, a risk-localized oracle bound replaces a discrepancy-dependent square-root penalty with an additional cost on the learning-risk scale. Label-level experiments use 29 recipes, a learned correction basis, estimated variance, and cost-matched holdout certificates. They identify both successful source correction and a shared-error failure mode. A matched diagnostic separates the probability-allocation cost of a general confidence event from its coefficient-envelope cost. The resulting framework provides computable risk comparisons for retaining calibration labels and choosing subsequent acquisition plans.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.