acceptodds
Under review as a conference paper at ICLR 2027

SMILE: Eliciting Expert Derivational Knowledge for Agent Benchmarks

Abstract

Evaluating AI agents on real-world work requires realistic inputs (prompts and references), deliverables, and expert-defined evaluation criteria (rubrics). Rubrics encode experts' criterial knowledge for assessing deliverables. These components alone do not make explicit the expert knowledge used to derive deliverables from inputs. We represent this derivational knowledge as atomic rules (ARs). Here, we introduce SMILE, a resource-efficient framework for building real-world agent benchmarks comprising inputs, outputs, and the derivation process connecting them. Expert-annotated or document-grounded ARs support evaluation of rule application in agent trajectories and expose relationships between tasks. These links support task maintenance: when a rule changes, its associated tasks can be identified and revised. They also support synthesis by recombining ARs into new tasks. As an instantiation of SMILE, we introduce SMILE, which comprises tasks from nine domains spanning Korean industry and public institutions. Our evaluation of frontier models reveals incomplete task execution and failures to satisfy critical requirements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.