ROAM: Mitigating Vision-Language Hallucinations with Calibrated Grounding Monitors
Abstract
Vision-language models can produce fluent responses that drift from visual evidence during generation. We introduce Read-Only Attention Monitoring (ROAM), a training-free decoding method that decouples visual-grounding monitoring from intervention. Offline, ROAM selects reusable monitoring heads using text-to-vision attention statistics and their stability across controlled image and text perturbations. The resulting monitor is calibrated once per model and reused across inputs. Online, the selected heads remain read-only: their image-attention entropy provides a grounding signal that adaptively controls the sampling temperature, sharpening the next-token distribution when the signal indicates weaker grounding. The intervention acts solely on the output logits, requiring a single forward pass per token without auxiliary decoding branches or direct modifications to internal representations. Analyses on LLaVA show that lower grounding scores are associated with higher hallucination rates before intervention, while ablations support the value of cross-environment calibration. Experiments across three VLM families and four benchmarks show the largest gains in long-form generation. On CHAIR, ROAM reduces sentence-level hallucination by 2.4–6.4 percentage points relative to vanilla decoding while increasing object recall on all three models, with near-vanilla end-to-end latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.