acceptodds
Under review as a conference paper at ICLR 2027

ShapingBench: Can Coding Agents Reproduce Mature Software Behavior Beyond Explicit Requirements?

Abstract

Coding agents are becoming increasingly capable of implementing software from high-level requirements, often expressed in natural language. However, existing benchmarks mainly evaluate behavior that is explicitly specified in the underlying task. We instead study how much of the behavior accumulated in mature software an agent can realize when those expectations are never individually stated. We introduce ShapingBench, a benchmark spanning 9 software domains, 45 mature open-source implementations, and 14,537 executable behavioral contracts. Behaviors are recovered from five independently developed implementations per domain and organized by how widely they recur across them. Behaviors shared by all five form the Shared Core, while the remaining behaviors form the Variable Surface. More than half of the Variable Surface appears in at least two reference implementations. We evaluate 4 LLMs across 3 agent frameworks, covering 10 model and framework combinations and 90 development trajectories. Behavioral coverage decreases consistently with cross-implementation recurrence, ranging from 62.1%-88.2% for behaviors shared by all five references to 12.4%-21.3% for behaviors observed in only one. Additional development improves coverage but yields diminishing gains, while many remaining failures involve behaviors repeatedly found across independent implementations. These results reveal a persistent gap between implementing explicit requirements and realizing the hidden behavioral expectations exhibited by mature software.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.