Generation and Judgment Diverge in Code Models: A Matched Benchmark with Derived Ground Truth
Abstract
Code-model benchmarks usually evaluate outputs whose correctness can be checked immediately, but many software decisions concern consequences that emerge only under future load. We ask whether models that can implement a change can also predict its system-level consequence. We introduce a matched benchmark built from synthetic service specifications with known resource decompositions. Each instance yields implementation tasks, graded by execution, and judgment tasks, graded against an operational capacity law validated on running services; four matched pairs probe the same underlying knowledge. Across six Qwen2.5-Coder scales from 0.5B to 32B parameters, chance-corrected implementation scores range from +0.275 to +0.616, whereas judgment scores range from -0.044 to +0.153, leaving a positive gap at every scale. Quantitative judgment remains at or below chance in all 24 scale-by-type cells, while categorical judgment improves with scale. A stripped arithmetic control shows that 1.5B and 3B models can solve the final argmax once the service representation is removed, localising much of the deficit to constructing the quantities required by the capacity law. Evaluation details also matter: raw accuracy can reverse model-family ordering, and a marker-strict answer parser can introduce a 45-point artifact. We derive the exact best-constant baseline for tolerance-scored numeric outputs, separate answer extraction from grading, and reproduce the main separation in a second model family.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.