Answer Pressure and Unsupported Answering in Language Models
Abstract
When instructed to answer only from a supplied document, a language model should abstain if the document lacks the required evidence. We examine whether removing an activation direction associated with answer pressure reduces unsupported answering while retaining supported responses. The design crosses supported and evidence-withheld contexts with neutral and pressure instructions and three intervention arms. In a paired analysis of author-reported model-run records, pressure increases unsupported answering by 15.0 and 13.5 percentage points in Gemma-2-9B-IT and Gemma-3-12B-IT. Target removal reduces these rates by 5.50 and 5.00 points relative to no intervention, reversing approximately 37% of the increase. Its advantage over assigned random controls is larger under pressure than under neutral instructions. Most of the net change is an increase in abstention; supported-accuracy point estimates decrease by 0.40 and 0.52 points. Comparisons with refusal and arousal controls and a separate ecological panel further characterize the effect. The contribution is a pressure-conditioned evaluation of selective responding using an established intervention operator. The instruction contrast jointly varies decisiveness and evidence-rule wording, and substantial unsupported answering remains. Appendix A documents data versions and provenance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.