acceptodds
Under review as a conference paper at ICLR 2027

TAKE THE HINT OR PUSH BACK? DECIDING WHEN A VISION-LANGUAGE MODEL SHOULD ACCEPT THE USER’S ANSWER

Abstract

Users of vision-language models (VLMs) often volunteer an answer. When it is wrong the model tends to adopt it, and inference-time interventions are built to suppress this sycophancy. When it is right, the same interventions discard part of the user’s help, LQCD nearly all of it. On a matched control that changes only whether the suggested answer is correct, Leading-Query Contrastive Decod- ing gains 52 points when an insistent user is wrong, loses 47 when the user is right, and keeps none of the 159 corrections the model had accepted from the user. Whether to take the hint can be decided without showing it to the model: in an opinion-free pass a correct hint ranks close behind the model’s own answer, a random wrong one far below. We propose accept-or-answer: the model answers the neutral question, and the user’s answer is returned instead when its logit mar- gin in that pass reaches a threshold calibrated on labelled queries. The policy needs one opinion-free pass, does not depend on how hard the user pushes, and can take its margin from another model. On 17 settings (seven benchmarks, three backbones, four phrasings, and hints written by people and by another model) it reaches 72.2% accuracy pooled over wrong and correct hints, against 69.3% for the best fixed policy of each setting, chosen after the fact; it is above that policy in 16 settings, significantly in 7 (4 after Holm correction), and significantly below it in none.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.