acceptodds
Under review as a conference paper at ICLR 2027

Intermediate Decoder Ladders: Beyond Temperature Calibration

Abstract

Learned readouts of intermediate language-model states are usually fitted to the final layer's full predictive distribution. We test whether this target still lowers the observed next-token negative log-likelihood (NLL) once a competing readout fitted to the teacher's argmax receives its own validation-selected temperature. Holding the intermediate states, frozen head, readout family and search budget fixed, we fit six decoder families on Qwen3-8B and Qwen3.5-9B and fix every fit and temperature before scoring 1,200 additional deduplicated documents. The full target lowers observed NLL by 0.124 and 0.224 nats, with simultaneous 95% intervals excluding zero and all 240 configurations agreeing in sign; per-arm temperature scaling removes 72% and 61% of the uncalibrated gap but not the remainder. A Llama-3.1-8B-Instruct/WikiText replication gives 0.266–0.271 nats across three training seeds, and its full affine readouts keep a 0.56–0.72-nat gain when rescored without refitting on 400 untouched articles, both in the first 128 tokens and in a later window. On Llama, reassigning whole teacher distributions among same-argmax positions removes 74–81% of the gain, and swapping only the tail while keeping each position's argmax probability removes 61–71%. The gain is largest for full affine maps, which stay 0.14–0.16 nats ahead of rank-256 maps under an extended budget, and the fitted readout beats the best calibrated cheap readout by 2.4–2.7 nats. For observed NLL, the full teacher distribution is the better of the two tested targets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.