ProdBench: Evaluating Requirement Recovery and Realization in Producer-oriented Software Development
Abstract
Real software requests leave crucial decisions unstated. We introduce PRODBENCH, a benchmark for evaluating an LLM Producer in agentic software development. A Producer turns an underspecified user request into an explicit implementation brief and oversees the development and validation of the resulting software. In our controlled protocol, the Producer gathers prior context, questions a fixed User Simulator, and directs a fixed Code Agent through implementation and revision. PRODBENCH contains 300 source-grounded Python tasks from 101 pseudonymized users, with 8,273 atomic requirements and 12,918 executable tests. Three matched request views vary the initial information while preserving each task’s target. Five metrics measure requirement recovery, delivery completeness, preservation of recovered intent, targeted clarification, and revision gains. Two annotators audit construction quality and validate recovery judgments, with agreement between the panel and each annotator of 91.9% and 90.8%. Across 13 Producers, delivery completeness exceeds a request-only coder using the same backend by 19.9–38.7 percentage points on single-file tasks and 10.1–22.3 on repository tasks. Yet richer context improves recovery far more than delivery. These results identify a persistent gap between recovering requirements and realizing them in software.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.