Accuracy-Optimal Stopping Can Reverse Treatment Effect Estimates from LLM Labels
Abstract
Repeated language model calls can improve the accuracy of outcome labels in randomized studies. We show that stopping optimized for pooled accuracy can nevertheless reverse the reported treatment effect while accuracy rises and the population and response kernel remain fixed. When stopping may depend on treatment, this occurs even with conditionally independent calls, answers better than chance on every record, and oracle posterior decoding. In a construction with at most three calls, we characterize the attainable region of accuracy and effect, identify an open budget interval where every accuracy maximizer reports the wrong sign, and derive the exact accuracy cost of preserving the effect. For any prescribed number of sign reversals, we construct one population and a finite call cap under which price optimal policies produce those reversals as call prices decrease and accuracy strictly increases. A first revision representation characterizes all attainable effect displacements within a call horizon, without requiring independent calls or Bayesian decoding. In the four type model, a known response catalog permits unbiased correction using at most two additional calls. Semisynthetic experiments using measured response catalogs and recorded model answers show that stopping optimized separately by arm produces substantial contrast displacement and severe confidence interval undercoverage on held out records. Ambiguous polarity pools can exhibit reversals, and policies with nearly identical accuracy can yield opposite effect signs. In the held out comparisons, fitted stopping rules shared by both arms produce much smaller mean absolute displacement with similar mean accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.