acceptodds
Under review as a conference paper at ICLR 2027

Learning to Be Fooled: pretraining grows knowledge but keeps the scar of framing susceptibility

Abstract

Language models learn from our text, misconceptions and all. Benchmarks cannot watch what they take from it: they score finished models, and when a question carries a cue, such as a leading phrasing or a suggested answer, the score mixes what the model knows with how far the cue moves it. When the score changes during training or after a fix, it cannot say which. We pair each item with a neutral form that keeps the fact and the answer options and drops the cue. The score then splits exactly into a neutral score and a paired gap , read per checkpoint, per item and per model component. On one pretraining run, OLMo-2-7B, four cues take four courses that no single score separates: a suggested answer (SycophancyEval) switches on within four billion tokens and never closes; a false premise (FalseQA) is resisted more with training; an irrelevant sentence (ARC-Challenge) costs a growing share of what the model knows; and a leading phrasing (TruthfulQA) opens a gap by 42 billion tokens that stays. TruthfulQA, the most criticised and the hardest case, is the only one whose cue-free form must be a rewrite: a third form that rewords the question but keeps the premise splits off a part due to the benchmark's adversarial wording. Its fall and partial return are then three processes, each on its own timescale: a short-answer prior leaves, the gap forms, and knowledge grows and makes most of the recovery. The score recovers, but the scar, six points of cue carried by a third of the items, stays through 3.9 trillion tokens, and three quarters of the whole gap survive post-training. All 32 base models we scored carry the gap; in the five also scored on the premise-kept form, a state-space model among them, the cue's own part is 3 to 7 points. The split reads fixes as trades: removing the cue's direction narrows the gap only by lowering the neutral score, and a consistency objective closes most of the gap by making the model heed any suggestion less, right or wrong. The instrument is released as a tool, twin, that builds the forms for any benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.