SELECTION REGRET: LOCAL ACCURACY MISRANKS CLOSED-LOOP RELIABILITY
Abstract
Local next-transition accuracy is a tempting way to choose components for a system that will eventually run in a loop. This paper shows that the metric can pick the wrong component once that system is evaluated end to end. In a controlled frozen-state routing setup that holds state access and termination fixed while varying transition scorers, teacher-forced edge accuracy is nearly identical across scorers—about 50–53%—but exact constrained-route accuracy differs by 36–88 points. Across 1,800 scorer-choice contexts, choosing by teacher-forced edge accuracy matches the closed-loop winner only 62.4% of the time. Mean selection regret is 9.84 points, compared with 13.37 for uniform-random selection; the 95thpercentile regret is 86.56 points, and 40.7% of contexts have tied top local scores. The mismatch persists when rank correlation is high: at L=4, correlation is 0.874 while selection agreement is 53.6%. Gold-prefix controls show that some failures precede generated-prefix feedback. Matched length-four continuation supervision brings L=4 accuracy from near zero to 75–82%; STOP-only data do not help, and margin-guided selection does not outperform random selection. On 90 AppWorld tasks and five backbones, a route-only metric selects a model with 0% task success under one official scaffold and 7.8% under a second; the best model reaches 13.3% and 23.3%, respectively. An offline metric that also checks call arguments selects that best model in both settings. Thus local next-step accuracy alone does not reliably select a component for closed-loop use. On the controlled surface, tied top local scores account for 64.2% of selection regret and can be detected without rollout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.