Agent-AutoBench: From Real-World Tasks to Executable Benchmarks
Abstract
Constructing agent benchmarks requires aligning task instructions, obtainable evidence, interaction rules, and scoring criteria while preserving information seeking and clarification. We present Agent-AutoBench, an automated method that transforms heterogeneous real-world materials into versioned executable task packages. Semantic hardening binds reasoning requirements to evidence and rubric criteria, and constrained rewriting updates these dependencies together. Search and fetch tools expose frozen sources, while a disclosure controller governs access to private user facts. Package verification and a tool-using Agent-as-Judge connect construction artifacts with execution feedback for task revision and capability-profile calibration. Daily-Life instantiates the method with a versioned release of 100 selected task packages spanning 11 domains. Archived multi-model evaluations characterize capability coverage and repeated-measurement variability. A historical, non-blind human audit records agreement with 96.47% of 7,815 criterion-level judge decisions. Complementary controlled tests demonstrate criterion-specific responses to output perturbations, restored arithmetic discrimination after a reference correction, and deterministic search and disclosure behavior across two implementations. Together, these results illustrate how evidence control, dependency-aware construction, and execution-grounded diagnosis support executable agent evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.