Probing System-Level Performance Representations in Code Models
Abstract
Can a code language model represent a system-level property more reliably than it reports that property in its output? We study capacity bottlenecks in rendered request handlers, where the correct stage is derived exactly from visit counts, per-visit costs, and the runtime sharing regime. Across six Qwen2.5-Coder scales from 0.5B to 32B on 1,200 instances, linear probes predict the bottleneck with 67.9–87.1% held-out accuracy, while surface-feature and shuffled-label controls remain near chance. A teacher-forced output read-out on the same held-out instances reaches 22.1–45.0%, leaving a 37.5–58.8 point gap. At 0.5B and 1.5B, the rank of the correct stage under the output logits is indistinguishable from chance; from 3B upward, the rank is significantly better than chance, indicating that the output distribution carries some bottleneck information even though it remains substantially less accurate than the probe. Layerwise profiles distinguish lexically explicit properties, which become decodable early, from composed properties, which emerge later; bottleneck decodability peaks around mid-depth and then decreases. The probe result also holds for DeepSeek-Coder 1.3B and StarCoder2 3B. Finally, steering with fitted probe directions is unstable across data splits: refitting changes the selected layer and reverses intervention effects. These results separate linear decodability, output accessibility, and intervention stability when probing system-level reasoning in code models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.