acceptodds
Under review as a conference paper at ICLR 2027

MOON: Reducing Object Hallucination with Repair-Aware Visual Guidance

Abstract

Vision-language models often describe objects that are absent from an image. External visual guidance can correct these errors, but it can also introduce false mentions or remove correct content. We introduce MOON, a controller that learns when and how to guide a frozen vision-language model. Its central idea is to predict what guidance will change. Adding a mention repairs an omission when the object is present and causes damage when it is absent; deleting a mention reverses these roles. MOON combines these effects with estimates of object presence and the unguided model's behavior to choose an action before generation. Across three backbone updates, it reduces average weighted object loss over 8–128-image paired-audit budgets by 1.9–2.3% relative to a direct risk predictor trained on the same observations. Both methods also use a pool of unguided target responses. The advantage extends to prompt and dataset changes and narrows as paired supervision increases. Controlled comparisons support repair/damage structure as a useful inductive bias for adapting visual guidance with limited paired data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.