The Specifications not Said: Can LLMs Turn Novice Developers' Wishes into Working Products?
Abstract
LLMs increasingly enable novice developers to build applications from scratch with natural language specifications, yet these developers often leave critical requirements unstated. Consequently, turning their immature specifications into working products not only requires implementing the application as instructed but recovering latent specifications as well. This gap is common in practice: in a deployed service where novice developers build mini-programs (a popular type of web application used by over one billion users) from natural-language specifications, only 3.94% of applications satisfy the developer after a single conversation. To evaluate this specification-recovery challenge with real-world developer data, we propose MiniProgramBench, a benchmark of 200 tasks across 10 application domains, built from developers' actual application creation and modification experiences and covering both frontend and backend functionality. MiniProgramBench pairs feedback from simulated developers with a frozen verification-point rubric, enabling evaluation of both specification recovery and product completion under a constrained feedback budget. Across 11 LLMs, verification-point pass rates reach at most 41.14% initially and 56.97% after up to two feedback-driven revisions. These revisions reveal that models struggle to incorporate feedback while preserving existing functionality. Controlled experiments further show that models perform substantially better with reference specifications than with self-inferred ones, underscoring specification recovery as a key bottleneck.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.