Validating Discovered Differences in Model Replacement Audits
Abstract
When an AI model is updated or replaced, we need to know how its behavior changes. Testing every possible input is impractical, so we search for behavioral differences between the original model and its replacement. But search creates two problems: a discovered difference may not hold on other inputs, while a search that finds no difference may still miss differences elsewhere. We separate what search discovers from what the evidence supports. Before search, we specify how discovered differences will be followed up on new inputs, and fresh validation determines whether there is enough evidence to support them. A false acceptance occurs when validation supports a difference that does not actually hold. Our validation controls this rate across search strategies. On Berkeley Function Calling Leaderboard (BFCL) tasks and controlled experiments, it validates differences on new inputs without false acceptances, while a simpler baseline produces a 49.44% false-acceptance rate. When search finds no difference, we instead analyze coverage of the untested inputs. Using a bound on how much the difference can change between nearby inputs, we determine when tested inputs can rule out differences elsewhere. Experiments show that limited coverage can leave existing differences undetected. Together, validation and coverage clarify what can and cannot be concluded about behavioral differences without exhaustive testing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.