acceptodds
Under review as a conference paper at ICLR 2027

Capability–Loss Dissociation in Language Models: A Theory and Empirical Study of Loss Signatures

Abstract

The standard measure of language-model progress is validation loss, but the ability of the validation loss to detect a capability depends on how often the capability is exercised in the evaluation corpus. This dependence is formalized by the loss signature, which is the product of a capability's token share and its per-token information gain. The signature quantifies the change in aggregate validation loss from acquiring a capability, and provides a criterion for when loss is an appropriate measurement tool. We assess this account on a synthetic dual-route language modeling task that separates parametric memorization from in-context retrieval. The predicted signature follows the observed loss gap across 160 controlled training runs with variants independently varying token share and information content. Even models with near identical validation loss can differ hugely in in-context accuracy, including a loss-matched pair with an 188x difference in ability. A gradient-level intervention with 76M parameters reproduces the dissociation, yielding a 69.5-point in-context accuracy gap between models evaluated at nearly identical aggregate loss (ΔL = 0.0027 nats). Direct measurement of the leakage term shows the capability’s own signature (fcΔLc = 0.1601 nats, 2.5% of operating loss) is almost exactly offset by a compensating gain on the competing memorization route (Λ = −0.1574 nats), leaving aggregate loss essentially unchanged. An ablation over five intervention weights shows this trade is smooth and monotone. On WikiText-103, induction is clearly visible in validation loss because its token share is large, illustrating that the framework is conditional rather than a blanket criticism of loss. We further derive a sample-size relationship showing that role-resolved evaluation can scale more favourably than aggregate loss for rare capabilities. Our results establish a practical principle: before using validation loss to measure a capability, estimate its token share and information content. When the resulting loss signature is small, or when a competing route can absorb it, capability-specific evaluation or role-resolved loss is required.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.