acceptodds
Under review as a conference paper at ICLR 2027

The Ruler Agrees With Itself, Not With the Task: Auditing Held-Out Loss as an Instrument for Context Selection

Abstract

When a language model's context is too long to keep whole, a selection rule decides which tokens stay. Such rules are often compared by held-out loss, which is cheap, label-free and precise. We ask whether the ordering it produces is the ordering a task produces. It is not, and where the two differ the loss is confident. We score twenty-one selection rules we constructed, each keeping half of a context's tokens, applied as token dropping on six frozen decoders and read on three loss corpora and four task benchmarks. Every ruler, a corpus or benchmark read with its own metric, reproduces its own ordering across halves of its data at a mean rank correlation of 0.956 to 0.996 corrected to full length; the loss corpora agree with one another at 0.83 on average, and a loss corpus with a benchmark at 0.44. Two instruments this reliable and this far apart measure different things. The disagreement sits at the top of the loss ranking: the loss prefers rules that keep recent or heavily attended tokens; the CBT benchmarks usually put a rule that keeps rare tokens first, LAMBADA usually second, and SQuAD's usual leader keeps the earliest tokens. It is not a matter of resolution: of the rule pairs that the WikiText-103 loss and a benchmark both order confidently, the benchmark reverses 30–44% (16–36% under the other corpora). Training a small selector on a frozen decoder against the loss does not remove the disagreement: on most task readings the selector stays in the rank band the loss gives it, but first place is decided by the task, and on CBT a rule the loss ranks in its bottom three usually beats the selector by point estimate. The obvious repair, a key-token gate that restricts the loss to tokens needing more than a short recent window, on average matches or modestly beats the plain loss only when that window sits well below the budget; tied to the budget, it measures the budget. We argue not for abandoning the loss but for auditing it before it chooses: the proxy's reliability, the target's, then their correlation read beside both.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.