acceptodds
Under review as a conference paper at ICLR 2027

CoDepth: Model-Free Generation of Compositional Coding Benchmarks

Abstract

High-quality coding problems are expensive to collect. Benchmarks such as SWE-bench rely on substantial human curation, while synthetic alternatives typically rely on costly model generation and validation. In this work, we show that new coding problems can be created at almost no cost by composing existing ones. We formulate problem generation as code composition: verified atomic problems are combined into level- tasks, and a task is kept only if every atomic problem remains unsolved and a joint reference solution passes the same check used at evaluation time. The entire task-construction process is programmatic and requires no language-model generation or validation. We instantiate this approach in four coding settings: multi-bug repair, proof completion, complexity analysis, and clone equivalence. The generated tasks are challenging for frontier models, which all degrade as compositional depth increases. The generated tasks are also useful as training data: rejection-sampling fine-tuning on them improves Qwen3.5-9B at every composition level.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.