Does Better Estimation Imply Better Selection? A Decision-Theoretic Audit of LLM Agent Updates
Abstract
Evaluating LLM-agent updates with low-cost proxies aims to reduce benchmarking expense, yet lower surrogate estimation error does not ensure better adoption decisions. We present a decision-theoretic audit framework that separates finite-pool mean squared error (MSE), same-pool decision regret, observed held-out gain, and acquisition cost. Our empirical benchmark evaluates three language models and twelve fixed update pairs on LiveCodeBench, using exact combinatorial integration across all target-sampling configurations to eliminate Monte Carlo acquisition noise over 12,384 recorded runs. In the post hoc subgroup of eight statically changed updates (Changed8), four-pair cross-fitting (CF4; two folds of two target task pairs) exhibits severe small-sample degeneracy at a 512-token proxy ceiling: both fitted slopes are zero in 95.71% of joint sampling configurations, yielding exactly the target-only mean on the same sample. Across all configurations in this regime, CF4 alters the corresponding target-only decision in only 0.076% of cases while requiring, on average, 2.40x as many cold-start acquisition completion tokens. At a 2,048-token proxy ceiling in the same subgroup, CF4 reduces mean MSE relative to unit-coefficient uniform correction but increases same-pool regret; externally calibrated correction reduces regret relative to CF4 but achieves lower observed held-out gain. Across all twelve updates, simultaneous 95% confidence bounds contain zero for both prespecified pooled utility contrasts. These findings reveal a low-budget regime where proxy adjustment adds substantial token overhead with little effect on decisions, demonstrating that surrogate evaluators must be benchmarked against matched target-only baselines through decision regret, held-out utility, and total acquisition cost rather than surrogate MSE alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.