acceptodds
Under review as a conference paper at ICLR 2027

Is It a Regression, or Just Noise? Calibrated Behavioral Regression Testing for Stochastic LLM Behavior

Abstract

An agent can fail a previously successful task even when its model and prompt are unchanged. Regression testing must separate this rerun variation from evidence that an update reduced reliability. We examine a repeated count test that compares each item's success counts with a one sided test and corrects across the suite for multiple comparisons. Our evaluation distinguishes two questions: unchanged version splits measure spurious alarms, while update runs measure agreement with labels marking large observed success rate drops. The test raised no observed alarms in unchanged version controls and selected many of the positive labels after updates. A structured function response study, repetition budget sweep and correction ablation show how the results depend on workload and sampling. Because the update labels share the tested runs, their agreement characterizes an empirical screen; independent reference runs remain necessary to measure detection of latent probability decreases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.