Knowing Without Saying: Code Models Represent the Bottleneck and Do Not Report It
Abstract
A code model identifies which stage of a request handler will saturate first and does not say so when asked. Over six scales of one code-model family, a 64× range, on 1,200 rendered services, a linear probe on the hidden states rises from 0.679 to 0.871 while the model’s own output read-out, obtained by teacher-forcing the answer line and ranking the candidate stages, rises only from 0.267 to 0.450 against a 0.304 majority; the gap is 0.375 to 0.588 at every scale and 0.421 at the largest. A saturated decision threshold does not explain it, and the mechanism changes with scale: at 0.5B and 1.5B the gold stage ranks at chance under the output read-out, so the property never reaches the output, while from 3B upward it ranks well above chance (p < 10−4) and the forced choice still recovers less than half of what the probe reads. Two further families, one spanning 25× by itself, reproduce both halves: the probe reads 0.846 to 0.867 and does not move with scale inside a family, while the read-out crosses from chance to above chance at the same size the first family crosses. What scale moves is depth, not accuracy: the property is read a third of the way into a sixty-layer network and is still not what the model says. The target makes this measurable without annotation. Which resource saturates is the argument maximising a product of a visit count and a per-visit cost, aggregated according to whether the runtime serializes its stages, so the labels follow from a capacity law and are exact and unlimited, and a logistic model on every surface feature reaches 24.6%. Companion targets that differ in how the information must be obtained turn one accuracy into a profile: the two lexically present properties are recovered at 100% by layer two at every scale, the composed one climbs with depth and scale, and the bottleneck peaks mid-network and declines, which separates a property a model reads from one it computes. Feeding the probe’s read-out back into the residual stream establishes a third result, one that concerns probing itself: the direction a linear probe returns is a sample of a subspace and does not survive a refit, inverting a p = 2 × 10−11 asymmetry when only the training partition changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.