Still Correct, Already at Risk: Pre-Onset Hallucination Forecasting in Vision–Language Models
Abstract
Existing approaches primarily predict hallucination risk in vision–language models (VLMs) before generation or detect local errors during decoding. We study the forecasting opportunity between these settings: predicting future hallucinations after generation has begun but while the output remains visually supported. We formalize this task as Pre-Onset Hallucination Forecasting: continuously predicting whether the remaining response will contain a hallucination using only currently available information. We find that still-supported generation provides predictive information beyond the generation-start state, enabling warnings several tokens before the first error, with gains that vary non-monotonically over the monitoring horizon. To study this predictive ability, we construct PreHallu, comprising approximately 85K generation trajectories and more than 180K unique continuations sampled from supported prefixes, enabling systematic study of hallucination onset and alternative futures. We further introduce RiskTrace, which combines current multimodal evidence with observed decoding history for continuous response-level early warning. Dynamic monitoring consistently outperforms generation-start prediction across InstructBLIP, LLaVA-NeXT, and Qwen3-VL. On InstructBLIP, using the same frozen scorer, dynamic monitoring raises response-level AUROC from 75.43% to 83.58%. Matched-false-alarm comparisons and lead-time analyses further show that pre-onset forecasting improves future-hallucination prediction while providing a window for verification or intervention before the first unsupported token appears.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.