acceptodds
Under review as a conference paper at ICLR 2027

PLAIT: From Hallucination Scores to Parent-Preserving Audit Plans

Abstract

Hallucination detectors estimate what may be wrong; human auditing must decide what is worth checking. Under a fixed review budget, these decisions are coupled: once a response and its evidence are open, additional claims become cheaper to inspect. We introduce PLAIT, which converts detector outputs into audit plans that account for these shared costs. A frozen detector, the parent, supplies initial values; generator hidden states learn corrections with a plan-regret objective; and an exact planner jointly selects response, claim, and span actions. A configurable constraint preserves a specified fraction of the parent plan's estimated value, and a separate verifier recomputes feasibility. Across four generators, claim-level PLAIT improves fixed-time recall by 3.54 percentage points over validation-selected exact score fusion. Gains persist after parent replacement and source-disjoint transfer. In a randomized study with 96 participants, PLAIT increases correctly identified unsupported characters per minute by 12.47% over HLA-RICO. Detection quality and audit utility are therefore distinct: useful scores must support good review decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.