Verifiability, Not Confidence: Cheap Multimodal Rumor Detection That Knows When to Refuse
Abstract
Multimodal misinformation detection faces a practical trade-off, in which discriminative detectors are cheap but less accurate while agentic pipelines are accurate but costly. We propose a minimal B vision-language classifier that answers each post in one forward pass, without an agent or retrieval. It approaches recently published detectors on AMG, FineFake, and Weibo at a measured USD per post. Auditing the residual errors reveals a deeper failure: many posts assert no verifiable claim, yet on content the detector never models its confidence is at its highest. We call this property *verifiability* and show that, under a fixed judge rubric, it is linearly decodable from frozen MLLM representations, with AUROC on AMG (in-domain), FineFake, and Weibo (both zero-shot transfers), and that the same readout transfers from English to Chinese and across model families down to B parameters while remaining near-orthogonal to confidence. A verifiability gate restores rejection where confidence fails, and a cheap resolver behind it keeps the cascade at or above direct accuracy. We believe verifiability offers a principled coordinate for abstention in multimodal detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.