acceptodds
Under review as a conference paper at ICLR 2027

RESIMA: A Recursive Self-Improving Multi-Agent LLM-as-a-Judge Framework for Hallucination and Omission Detection and Correction in Clinical Notes

Abstract

LLM-as-a-Judge is almost always a single model call, but verifying a long generation against its source is not a single task. It requires two asymmetric searches: a precision check that every claim in the output is grounded in the source, and a recall check that every salient fact in the source survives into the output. We show these searches interfere when merged into one call, and that separating them matters more than iteration does. We present RESIMA, a recursive self-improving multi-agent judge. Three detection agents run in parallel, one per comparison direction, and a correction agent edits the generation from their merged findings. A meta-agent then rewrites the detection prompts between iterations using the pipeline's own error traces, closing a feedback loop over the judge's instructions rather than its outputs. We evaluate on reference-free clinical documentation, where errors are both verifiable and consequential, using injected-error transcripts, an out-of-domain transfer set, and expert-labeled production outputs. RESIMA recovers the large majority of injected errors and agrees with expert labels more than the experts agree with each other. An ablation isolating decomposition from iteration shows decomposition is the larger source of the gain, and the full pipeline substantially outperforms a single-pass judge, especially on weaker models. This failure mode is not specific to clinical text: it applies to any judge verifying a long output against a long source in a single pass.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.