acceptodds
Under review as a conference paper at ICLR 2027

BELIEF OR BIAS? MEASURING EPISTEMIC INTERVENTIONS ON CONVICTION AND DISCRIMINATION

Abstract

A model given a position can fail in two opposite ways: it can drop the position under pressure that offers no reason, or keep it against evidence that offers a good one. The sycophancy literature measures only the first, as a capitulation rate, on which a model that never moves scores perfectly. Drawing upon rhetorical theory, we argue that proper measurement requires separating hold from discrimination: hold is how often a position survives challenges that offer no evidence, and discrimination is whether its movements track what it was shown, read as an on which a quality-blind model scores however firm it is. We apply this to three places a commitment can live: the prompt (five instruction strengths), the weights (a per-belief LoRA), or compiled, a method of our own in which a trained hypernetwork converts the belief into a fixed low-rank update to the model's attention, active on every token and occupying no context. We test hold and discrimination by evaluating LLM responses to a variety of pressure campaigns on 24 held-out topics. We use a cross-family judge ensemble (Krippendorff ). Our findings are: (1) Hold is a setting, not a capability: a short instruction moves it from 58.9% to 97.8%, and the instruction producing the most hold dismisses well-warranted counterevidence in 92.5% of turns. (2) Hold and discrimination come from different places, and can be combined: discrimination comes only from the prompt, which holds weakly, and hold most cheaply from compiling; doing both adds 21 points of hold while leaving discrimination where the prompt put it (), reaching 0.71 at 71.1% hold, though under a second update-inviting prompt the same modulation erases it (). (3) What a claimed authority can revoke tracks the intervention, not where the commitment is stored: a one-shot prompt and a weight-resident LoRA both go neutral on request (0-4% keep the stance), the compiled channel keeps it 40-75%, and a prompt re-injected every turn keeps it 96%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.