acceptodds
Under review as a conference paper at ICLR 2027

Residual Scaling for Language Models: An Exact Layer Score and an Error Budget for Composing It

Abstract

For scalar gains on the residual contributions of a frozen language model, the signed gradient–residual product is the exact loss derivative at identity. We locate where its predicted benefit is lost when composing multiple layers. The derivative ranks single-layer amplification at Spearman – on matched prompts and – on held-out prompts, whereas measured ablation ranks it at to on the same Qwen3.5-9B items. The gap between summed predictions and joint gain splits exactly into cross-item transfer, a single-coordinate finite step and cross-layer interaction. Across four checkpoints of three families, the finite step dominates on GSM8K at ; for three-layer profiles, interaction is at most 30% of it. Transfer grows linearly and the finite step locally quadratically with strength. Coefficients frozen at predict switches between these terms on ARC-Challenge: registered fp16 measurements on the original questions locate them at –, and on disjoint questions at –, within the registered tolerance. Choosing triples by measured single-layer gains lowers joint-loss regret from to ; measuring eight derivative-ranked candidate blocks recovers this choice at an estimated 45% of the recorded single-block stage time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.