When Honest Evaluation Fails to Compose: Learner-Protected Targeting Curves
Abstract
Evaluating each learned treatment rule on fresh data does not ensure valid inference when the evaluations are combined. Earlier evaluation noise can influence later model selection and change the variance of later errors, invalidating a bootstrap that freezes the learned path. A two-candidate characterization identifies when this failure persists and when frozen inference remains valid. We construct asymptotically valid confidence bounds for the value of the realized rules, simultaneous over treatment budgets and version weights. Versionwise bands combine rectangularly even with overlapping evaluation samples. For disjoint evaluations with Gaussian limits, studentization gives radial bounds with a sharp common critical value at each budget. A matched four-route experiment separates the benefits of data reuse from those of calibration geometry: using all later cohorts can improve selected utility, while radial bounds improve precision for balanced mixtures. A matched selection–decision simulation shows a concentrated benefit of protection. At one predeclared cost, it leaves 99.6% of decisions unchanged and improves average utility by preventing harmful deployment in the remaining 0.4%; at a lower cost, it forgoes useful treatment. On Hillstrom, reuse increases the certified net value, while exact-budget targeting advantages remain uncertified. A finite-batch correction supports deployment with integer treatment quotas.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.