acceptodds
Under review as a conference paper at ICLR 2027

LLM Judges Verify Presence, Not Absence: Omission Blindness in Generated Clinical Notes, Its Mechanism, and What Recovers It

Abstract

LLM judges score generated text at scale, and the benchmarks that measure them ask whether what a text says is correct, not whether what it should say is there. We measure that second capability where it is consequential and, unusually, constructible: ambient AI scribes writing clinical notes, where published human audits find omission the dominant error class. Public corpora cannot supply the answer key - every clinician reference note we audited is materially discrepant with its own transcript - so we release 500 single-error note pairs built from transcript-derived, audited fact sheets: 298 in which a named fact is certainly absent, graded by severity and by surviving trace, against 202 added-or-altered controls. Across an eight-design ablation, paired discrimination (0.5 is a coin flip) reads 0.79-0.94 on added or altered content against 0.50-0.63 on omissions, and on single notes every design flags notes with omissions about as often as perfect ones. Criterion scope, output format, voting, wording and automated prompt optimisation all move the operating point without creating detection, each with a bound on what we could have missed, and a larger reasoning budget lifts discrimination without closing the asymmetry. What recovers detection is restructuring the task - list the facts the source establishes, then verify each against the text as a closed presence check. Two methods reach that restructuring independently, and a control separates its parts, the fact list carrying about a third of the recovery and closed per-fact verdicts the rest. Held out, reading those verdicts as a rule rather than averaging them into a score is worth 2.7 times the detection at an identical false-alarm rate, and transfers across judge families near-unchanged, while no threshold set on constructed pairs survives real scribe output. Omissions whose restatement survives elsewhere defeat both. A physician author validated 70 items in a sitting covering six stages, three of them structurally blinded, and sided with the pipeline on all 10 notes where it and the best monolithic judge reach opposite conclusions; the severity rubric has since been graded blind by an independent clinician - not an author - agreeing to within a grade across the companion census's two clinician sittings. We release the benchmark, the prompts and every judgement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.