Knowing the Parts Is Not Enough: Measuring Composition After Matched Precursor Success
Abstract
Aggregate code-generation benchmarks report how often a model solves a problem. They cannot separate two failures: the model never demonstrated the required capability, or it succeeded on a matched precursor task for each constituent yet could not combine them. ConceptBench targets the second case by conditioning composed-task failure on matched precursor success. Each of 74 families over five concept pairs is instantiated as three tasks sharing one semantic core — capability A alone, B alone, and their composition AB — where A and B are projections of the AB rule onto input domains in which the other capability is structurally absent. The outcome is the raw PPF (both precursors pass, the composition fails), reported as P(AB fails | A ∧ B) with its eligible-family denominator: a protocol-conditional diagnostic, not an estimate of how often composition fails. Under one pinned inference protocol, models differ sharply even on the same jointly eligible families. On the same ten families of one concept pair, for which both models passed both matched precursors, Qwen2.5-Coder 14B produces nine raw PPFs and Qwen2.5-Coder 32B none; two of the nine are mechanically verified precursor-coverage gaps, and setting those observations aside the contrast remains 7–0. Eight further matched common-support cells differ by at least three raw PPFs, in both directions. Eligibility selection alone therefore cannot explain this heterogeneity, though the contrast is not shown to be composition-specific. Pair-level rates also vary, but their rank ordering is identical to that of unconditional AB failure (P3 > P2 > P5 > P1 > P4), so pair ordering alone does not identify composition-specific structure. Mechanical screens expose specific confounds; beyond them, mechanistic attribution was not human-adjudicated, and no global count of composition-specific failures is reported. As descriptive external validation, two frontier models with reasoning disabled each produce one raw PPF; with reasoning enabled and a larger output budget, both failures become passes and each produces one new raw PPF.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.