ShapingBench: Can Coding Agents Reproduce Mature Software Behavior Beyond Explicit Requirements?
Abstract
Coding agents are becoming increasingly capable of implementing software from high-level requirements, often expressed in natural language. However, existing benchmarks mainly evaluate behavior that is explicitly specified in the underlying task. We instead study how much of the behavior accumulated in mature software an agent can realize when those expectations are never individually stated. We introduce ShapingBench, a benchmark spanning 9 software domains, 45 mature open-source implementations, and 14,537 executable behavioral contracts. Behaviors are recovered from five independently developed implementations per domain and organized by how widely they recur across them. Behaviors shared by all five form the Shared Core, while the remaining behaviors form the Variable Surface. More than half of the Variable Surface appears in at least two reference implementations. We evaluate 4 LLMs across 3 agent frameworks, covering 10 model and framework combinations and 90 development trajectories. Behavioral coverage decreases consistently with cross-implementation recurrence, ranging from 62.1%-88.2% for behaviors shared by all five references to 12.4%-21.3% for behaviors observed in only one. Additional development improves coverage but yields diminishing gains, while many remaining failures involve behaviors repeatedly found across independent implementations. These results reveal a persistent gap between implementing explicit requirements and realizing the hidden behavioral expectations exhibited by mature software.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.