Is It a Regression, or Just Noise? Calibrated Behavioral Regression Testing for Stochastic LLM Behavior
Abstract
An agent can fail a previously successful task even when its model and prompt are unchanged. Regression testing must separate this rerun variation from evidence that an update reduced reliability. We examine a repeated count test that compares each item's success counts with a one sided test and corrects across the suite for multiple comparisons. Our evaluation distinguishes two questions: unchanged version splits measure spurious alarms, while update runs measure agreement with labels marking large observed success rate drops. The test raised no observed alarms in unchanged version controls and selected many of the positive labels after updates. A structured function response study, repetition budget sweep and correction ablation show how the results depend on workload and sampling. Because the update labels share the tested runs, their agreement characterizes an empirical screen; independent reference runs remain necessary to measure detection of latent probability decreases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.