acceptodds
Under review as a conference paper at ICLR 2027

HalluClear: A Diagnostic and Evaluation Suite with Mitigation Probe for Hallucinations in GUI Agents

Abstract

As GUI agents increasingly benefit from industrial-scale training, ungrounded hallucinations can still lead to cascading failures during sequential GUI interaction. Unlike hallucinations in general vision-language tasks, GUI-agent hallucinations span both visual perception and task reasoning and are further complicated by partial observability. Existing benchmarks primarily assess task success, visual grounding, or action fidelity, but lack fine-grained hallucination diagnosis and reliable evaluation. To bridge this gap, we introduce **HalluClear**, a suite for diagnosing and evaluating hallucinations in GUI agents, with a mitigation probe. HalluClear consists of three components: (1) a diagnostic framework that formalizes GUI-agent hallucination as unsupported belief-action commitment under partial observability and derives a fine-grained taxonomy from empirical failure analysis; (2) a three-stage evaluation workflow that combines an expert-annotated benchmark with qualified VLM judges for credibility-calibrated hallucination measurement; and (3) a lightweight mitigation probe that instantiates a structured post-training intervention motivated by patterns identified through our diagnostic analysis. Experiments with representative GUI agents on public benchmarks reveal systematic hallucination patterns across agent families and failure subtypes, providing diagnostic information complementary to conventional visual grounding and action-fidelity metrics. Across three open-source GUI agents, a post-training intervention using 9K samples consistently reduces hallucination rates while improving visual grounding and action fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.