LIFE: LOCAL INTRINSIC DIMENSION-GUIDED FEATURE EDITING FOR ONE-TOKEN, ONE-LAYER HALLUCINATION REPAIR IN VISION-LANGUAGE MODELS
Abstract
Large vision-language models (LVLMs) can hallucinate, producing fluent responses with details that the image does not support. Hallucination leaves traces in the model’s internal representations, but the decoder layer where these traces are clearest varies across models, benchmarks, and individual inputs, and existing methods inspect or edit a layer that is fixed in advance. We introduce **LIFE** , ocal ntrinsic Dimension-Guided eature diting, an inference-time framework that decides where to inspect, whether to intervene, and how little to change. **LIFE-Detect** uses intrinsic dimension to locate hallucination-sensitive layers without labels, at two scales: the population-level intrinsic dimension of calibration representations identifies a shared layer for each model and benchmark, and the local intrinsic dimension (LID) of an individual input selects a layer for that input. A small MLP probe at the selected layer decides whether the response will likely be a hallucinated one. **LIFE-Mitigate** then edits one hidden state, the representation of a single token at a single decoder layer, by moving it along a direction given by the probe with a step size chosen for that input, and regenerates the response. The LVLM is never fine-tuned and never back-propagated through; the only trained component is the probe, and the correction optimizes one scalar. Across three LVLM backbones and five benchmarks, LIFE reduces the AMBER hallucination rate from 21.3% to 18.8% on LLaVA-1.6 and from 14.2% to 11.0% on Qwen3-VL, and more than halves sentence-level CHAIR on all three backbones, at a per-sample cost dominated by regenerating the response.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.