The Decision Value of Imperfect Provenance: A Minimax Characterization for Shadow Evaluation
Abstract
Online systems use shadow evaluations to compare candidate models or agents with the production choice. Same-request comparisons can cancel shared variation, but imperfect pairing leaves their quality uncertain. We study how provenance determines the value of a fixed shadow budget under bounded i.i.d. losses and marginal-preserving independent resampling. Two channels can have the same erased-score law, matching accuracy, and mutual information about request matching, yet yield constant versus square-root minimax regret at full budget. The relevant quantity is the posterior-contamination spectrum , where is the probability of independent resampling conditional on the observed provenance. This spectrum governs achievable comparison precision and the hard-family information bound for complete adaptive transcripts, giving minimax regret bounds that match up to logarithmic factors. It leads to calibrated inverse-variance weighting and an order-only method that retains the polynomial regret rates without calibration. In controlled experiments at , full budget, and a near-clean tail exponent of , the maximum tested-gap mean regret is for calibrated weighting, for unweighted pairing, and for order-only service.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.