acceptodds
Under review as a conference paper at ICLR 2027

EvalScientist: Automated Benchmarking and Hypothesis-Driven Diagnosis for Frontier LLMs

Abstract

As frontier large language models (LLMs) advance, they need to be evaluated on more capabilities and in more complex settings, but expert-curated benchmarks are costly to build and quickly saturate. Automated methods reduce this cost, but they are largely limited to what available data or domain-specific formats support, and they show where models fail without diagnosing why. We introduce **EvalScientist**, an automated framework that handles natural-language evaluation requests end to end. It designs tasks using research on external sources, constructs them in a unified task representation that covers all task types, and evaluates target models on them. It then performs hypothesis-driven diagnosis, using targeted probes to test candidate weaknesses that may explain observed task failures. On 8 conventional requests covering widely benchmarked capabilities such as knowledge, reasoning, and coding, EvalScientist outperforms existing automated frameworks and GPT-6-Astra in correctness, alignment with the request, diversity, and difficulty. We further identify 15 new evaluation requests from authoritative sources such as AI safety policies of governments and frontier AI developers. These requests target frontier agentic and safety-related behaviors that no existing benchmark fully covers. EvalScientist constructs correct and challenging benchmarks for all of them, whereas existing automated frameworks either fail to accommodate these requests or generate tasks that are incorrect, misaligned with request, or overly easy. In blind pairwise human evaluation, judges prefer EvalScientist's tasks in 62.5–87.5% of comparisons against other automated frameworks and in 55–60% against expert-curated benchmarks. Hypothesis-driven diagnosis yields findings supported by correct and aligned tasks, whereas simply generating more tasks like those the target models fail drifts toward ambiguous tasks, such that low scores increasingly reflect task defects rather than model weaknesses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.