They Know When They're Unsure: Confidence-Gated Acceptance as a Two-Sided Defense against Sycophancy in Video Language Models
Abstract
Video large language models (Video-LLMs) can answer visually grounded questions correctly, yet systematically reverse those correct answers when a user pushes back with a confident but wrong claim—a multimodal form of sycophancy. We first show that this failure must be measured two-sidedly: a defense should hold correct answers (should-hold) while still correcting genuine mistakes (should-flip). Under a two-sided balanced-accuracy metric, six language-level defenses—including self-consistency, chain-of-thought, an explicit "be confident" instruction, and our own structured visual-anchoring prompt—all collapse to the trivial never-change baseline (), because the model cannot express its own uncertainty in language. Yet the internal signal exists: first-answer probability predicts first-answer correctness with AUROC – across all models. We exploit this with Confidence-Gated Acceptance (CGA), a training-free defense that accepts a user's claim only when the model's own answer probability is low, using a single extra forward pass and one cross-validated threshold. Across five video-QA benchmarks and multiple open models, CGA significantly beats naive acceptance and the never-change baseline in all eleven evaluated cells, with the largest gain (pt) on the poorly-calibrated Qwen3-VL-8B. A threshold learned on Video-MME also transfers to TempCompass without retuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.