acceptodds
Under review as a conference paper at ICLR 2027

NL2FullStack: Evaluating Data Modeling and Full-Stack Application Generation from Natural Language

Abstract

Large language model (LLM) coding agents can build complete database-backed applications from written briefs, but existing benchmarks grade user-visible workflows or the resemblance of a generated schema to a single reference, so backend business logic and data integrity are overlooked. We present NL2FullStack, a benchmark for data modeling and full-stack application generation from natural language. An agent receives a business brief and a binding engineering contract, then must ship a bootable application with a backend that implements the required workflows and a database that enforces the specified rules. Grading is deterministic and uses no LLM judge: the harness replays hidden, seeded workloads over the application's declared routes to test business logic, probes the database directly to verify integrity constraints, and checks that the interface displays backend rows. Across 89 tasks with five attempts each, Claude Opus 4.8 passes 62.9% of attempts, ahead of GPT-5.6 Sol (48.8%), Kimi K3 (44.5%), DeepSeek V4 Pro (29.2%), and Qwen3.8-Max (27.2%). The top four models pass the database tier on 87–96% of attempts, so nearly all the gap between models comes from the backend tier, where pass rates fall from 68% to 38%. For every model, the dominant failure is rejecting valid input under constraints absent from the contract. To mitigate contamination from public benchmarks, we release a nine-task public set with the grader, agent harness, and every graded attempt, and hold out eighty tasks as a private evaluation set.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.