acceptodds
Under review as a conference paper at ICLR 2027

Measure First: The Noise Floor of an Autonomous Research Loop Dominates Its Accept Rule

Abstract

Autonomous research agents change their code, retrain, evaluate, and keep the change if the score went up. Recent work proposes better statistical tests for this keep-or-discard decision. We argue that the first problem is that nobody measures how noisy the score is, and we measure it in four ways. First, we re-run one production LLM distillation-and-quantization pipeline, unchanged, thirty times in each of sixteen settings (480 runs): run-to-run spread varies more than 13-fold across settings, and the improvement the pipeline's own authors were chasing cannot be distinguished from this noise (its 95% upper bound is 0.54 of it). Second, we read the source code of 26 systems: sixteen of eighteen deployed agent loops keep a change after a single run, and none lets a noise estimate reach the decision. Third, we replay eleven decision rules on the measured runs. The anytime-valid e-process, written against a noisy baseline as is natural, is not valid (false-accept 0.038, 0.148 and 0.278 at 8, 32 and 128 runs); corrected, its best variant reaches 0.048 power at half a standard deviation, where a five-run Welch t-test reaches 0.170 and needs no external noise estimate. Taking that estimate from three runs breaks the corrected rule in all sixteen settings. Fourth, we evaluate the same frozen weights 3,627 times on seven NVIDIA GPU models. Scores move in whole questions; 13 of 39 model, task and GPU combinations return the identical score every time while others drift by up to thirty questions; two GPUs that are each perfectly repeatable still disagree by twenty; and running inputs one at a time instead of batched removed the drift in every case we tested. We recommend two things: before trusting any comparison, re-run the unchanged pipeline several times on the hardware the loop will use, to learn how much the score moves on its own; and decide with a test that estimates the noise from the runs themselves, such as a Welch t-test, rather than one handed a noise level from outside.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.