acceptodds
Under review as a conference paper at ICLR 2027

Scaling Where It Matters: Predictive Contraction for Efficient Internal Scaling in LMs

Abstract

Scaling language models with larger networks or more inference trajectories improves performance but significantly increases computational load. This raises an alternative: can extra computation be allocated selectively inside the Transformer, rather than globally, to improve performance with modest overhead? Our core idea is that prediction-relevant refinement is highly non-uniform across layers. Using a fixed logit lens, we observe sharp predictive entropy contraction at a few layers and treat it as an observable signal of strong prediction-relevant refinement. We call such layers predictive concentration layers. Based on this signal, EntroBranch performs efficient localized scaling by selecting the layer and branch width from frozen-backbone calibration, adding prefix-conditioned branches there, and aggregating their outputs with a token-wise gate before resuming single-path computation. Across Qwen3 and Gemma4 on code, mathematics, and multiple-choice tasks, EntroBranch improves overall scores by 33.70% on average over single-path prefix tuning with only 4.58% additional per-token inference FLOPs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.