acceptodds
Under review as a conference paper at ICLR 2027

AppSci-Bench: Can LLMs Build AI Agents and Make Them Better?

Abstract

Large language models (LLMs) are increasingly used not only to power AI agents but also to build them. While many benchmarks assess the quality of LLM-powered agents, few ascertain how good LLMs are at _building_ them. To this end, we introduce AppSci-Bench, where we construct a setting akin to how human **App**lied **Sci**entists build agents for businesses - through eliciting task requirements from stakeholders, exploring the business's infrastructure, and scientifically improve the target agent by designing experiments and internal benchmarks. AppSci-Bench features seven realistic domains - five new ones (Slides Generation, Order-to-Cash, Compliance Review, Customer Call Insights, and CRM Insights) and two domains recasted from -bench (Airline and Banking) - each with a held-out evaluation suite customized for that task. We evaluate seven frontier models and find AppSci-Bench challenging: the best model, Claude Opus 5.0, only reaches a mean-minimum score of **59.1%** across 3 build attempts; other models do significantly worse. Models struggle with eliciting task requirements (only 64% for Opus 5.0), build representative internal evaluations (models saturate their own evaluation 65% of the time but fail to solve the held-out task), presenting challenges for future research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.