acceptodds
Under review as a conference paper at ICLR 2027

Beyond Static Benchmarks: An Active Testing Paradigm for Large Language Models

Abstract

Static benchmarks remain the dominant way to evaluate large language models, yet fixed questions are vulnerable to saturation, contamination, and overfitting. More fundamentally, their test distributions do not adapt to the model under evaluation, potentially spending evaluation budget on regions that are already well understood or largely solved. We propose Active Testing, an evaluation paradigm in which observed model behavior dynamically guides where to test next, with emphasis on regions that are both failure-prone and uncertain. To demonstrate its effectiveness, we instantiate this paradigm for web-search agents by turning Wikidata into a large, structured, and verifiable task space with 12 controllable dimensions. We realize the active tester with a random-forest surrogate, enabling sample-efficient search and targeted discovery of weak regions in a large discrete task space. Across six agents, accuracy drops by an average of 28.0 percentage points from the initial random baseline to the final attack round. Rankings by Active-Testing AUC also diverge from initial-accuracy rankings, revealing performance differences that the shared starting distribution obscures. The tester converges to different regions for different models, exposing model-specific weaknesses while also uncovering shared failure patterns across agents. These results motivate a shift from static benchmarks toward active testing protocols, offering a more proactive and diagnostically informative approach to evaluating frontier models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.