Measured Behavior, Recorded Zero: Measurement Failures in Self-Improving Coding Agents
Abstract
Self-improving coding agents evolve a scaffold around a frozen language model and select variants using benchmark performance. We study whether that selection signal reliably reflects what the agent and the search actually did. On the 59-task Polyglot subset used by the Huxley–Gödel Machine (HGM), the default self-hosted setup records zero tool use at 7B, 14B, and 32B, yet the 14B transcripts contain dispatchable tool calls on 30/59 tasks. Recovering those calls changes tool use from 0/59 to 32/59 while resolve rate remains 0/59. In two evolutionary runs, all 14 proposed children are invalid Python and are archived with the same utility (0.0) that a valid child solving no tasks would receive. Replacing the proposal model with a frontier hosted model changes validity from 0/14 to 14/14 while preserving the same destructive edit pattern. Across ten configurations, resolve rates remain between 0/59 and 4/59. These results do not show that scaffold evolution fails. They show that, in this deployment regime, the recorded objective can collapse distinct behavioral states and can therefore obscure whether a search step acted, failed, or produced an executable candidate. We provide a network-level measurement layer and a set of inexpensive controls for evaluating self-improving agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.