When Does Readout Geometry Explain Sycophantic Answer Flips?
Abstract
Expert advice can overturn a language model's correct answer. We study when directions derived from answer-token output weights also support effective internal interventions. Under root mean square normalization (RMSNorm), we derive the exact finite change in an answer-pair margin. Across three models and two tasks, near-output removal along the projected raw output-weight difference yields 64.4% to 96.6% correct selection on initially incorrect prompted answers, versus 0.5% to 2.8% for norm-matched random removal. We then compare the direction adjusted for normalization gains with the downstream gradient. At relative norm 0.001 in 32-bit arithmetic, its mean response ratio to an equal-norm gradient intervention is only 0.4% to 8.1% at the tested middle layers, reaching approximately 100% at the final layer. The exact expression reduces margin-prediction error by 82.3% across 28 prompt conditions. Precision controls preserve the layer effect, and a factorial experiment separates authority from certainty. Readout geometry thus supports effective control near the output; earlier interventions require directions informed by the remaining computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.