Generate, Verify, Retract: Exploiting the Generation–Verification Gap in Vision-Language Models
Abstract
Vision-language models (VLMs) can generate visually unsupported objects, which then become part of the autoregressive context and may influence subsequent generation. We uncover a substantial **generation–verification gap**: on paired generated object events, the same frozen VLM's explicit object-existence judgments detect hallucinations substantially better than generation-token probabilities (AUROC 0.8584–0.9022 vs. 0.6093–0.6155). This suggests that although a VLM may fail to generate only visually grounded content, it can often recognize an unsupported object once that object has been proposed. Motivated by this observation, we introduce **RETRACT**, a decoding framework that transfers the model's explicit verification capability into its generation process. RETRACT distills object-existence judgments from the frozen VLM into a lightweight verifier operating on candidate-specific internal representations. During decoding, each object candidate is verified after being processed but before being retained in the generation context, allowing unsupported candidates to be locally retracted before they influence subsequent generation. This design avoids repeated explicit verification queries and requires no modification to the VLM parameters. Across multiple VLM backbones and hallucination benchmarks, RETRACT consistently reduces object hallucinations while preserving generation quality, with only modest inference overhead. Code is available at [https://anonymous.4open.science/r/RETRACT/](https://anonymous.4open.science/r/RETRACT/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.