acceptodds
Under review as a conference paper at ICLR 2027

You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

Abstract

A frozen language model under-uses the evidence in its hidden states and answers confidently when the context contains no answer. We fix both on the residual stream without changing any weight. For accuracy, a small trained network (the probe) adds an input-dependent vector to the residual at several middle layers; a removable hook, it raises Qwen2.5-1.5B on αNLI (abductive reasoning) from 51% to 73% accuracy. For abstention, a fixed vector (the direction), the mean hidden state of contexts that contain the answer minus that of contexts that do not, scores the hidden state; a low score abstains. On four QA datasets with human answerability labels (Qwen2.5-7B), moving the direction to a dataset it was not fit on costs under 3 percentage points of decision accuracy (from 0.83 in domain) on average and still beats asking the model whether its context suffices; the same judgment trained as a classifier into the probe's activation edit or into LoRA weights loses 16 and 22 points and ends below asking. We set out to put both functions on a single forward pass, and found that the probe's edit changes the very hidden state the direction scores, degrading it. A small map — trained without labels, only to reconstruct the unedited hidden state from the edited one — removes 63% of that interference on 1.5B. Refit on the hidden-state pairs of LoFiT or LoReFT, the same recipe repairs those methods too: the steering module is replaceable. The resulting system, YOPO, answers and abstains in one forward pass. On αNLI it raises three-way accuracy (correct answer when the context suffices, abstention when it does not) from 0.38 to about 0.75 on 1.5B, and stays within 0.02 of a two-pass reference, which scores a separate probe-off pass, on 1.5B, 3B and 7B, at half the cost. We also observe improvements on the four QA datasets, where answerability is human-labeled and answers are generated as free-form text. Near-parity with the two-pass reference holds across nine backbones spanning six model families without backbone-specific hyperparameter tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.