acceptodds
Under review as a conference paper at ICLR 2027

Complexity Ladder: A Controlled Test of Specification Complexity in Code Generation

Abstract

A previous study, ComplexityKink, found on a scraped corpus of about 5,000 prompts that the structural complexity of a programming prompt predicts whether a language model solves it, and that the relationship has a breakpoint. An observational corpus cannot say whether complexity itself did that, because prompts that are complex in the wild differ in many other ways at once. This study controls the prompt instead of sampling it. We take one hundred base problems, ninety-eight of them contest problems taken as written, and produce five further versions of each by adding one requirement at a time, holding the base task, the function signature and the test harness fixed: 600 prompts, each rated on the same frozen rubric by four judges. Sixteen models generate every prompt five times, 48,000 draws in all. For each model separately, a pre-registered and calibrated procedure tests whether family-adjusted correctness departs from a single straight line along the rated complexity composite and, where it does, locates the change and reports its direction and size. No model meets the registered criterion for a breakdown: one, GPT-5.6 Sol, rejects the straight line but its departure cannot be located on the grid, and the other fifteen do not reject at this design's precision. Every model is nonetheless lower at rung 5 than at rung 0, by 2.5 to 40.8 points of graded correctness. A descriptive split, chosen after the result, shows that for six models most of that decline arrives as answers cut off at the output limit, which score zero, while correctness conditional on a program running changes little, and that for the others the programs that run get worse. These results do not establish that breakpoints are absent, and they do not explain the earlier observational finding; they show what a controlled ladder, read with a calibrated procedure, can and cannot see.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.
Complexity Ladder: A Controlled Test of Specification Complexity in Code Generation | acceptodds