acceptodds
Under review as a conference paper at ICLR 2027

RAV: Retrieval-augmented Training of LLM Judges for Chest X-ray Report Verification

Abstract

The use of LLM-as-a-judge model is popular for evaluating language generation in radiology report generation. However, reference-free evaluation is prone to self-preference bias preferring hallucinated plausibility to clinical factuality when no objective ground truth is present. In this paper, we take a new approach to fact-checking by training such judge models to exploit the consensus present in prior radiology reports of similar patient images. Specifically, a new loss function is designed to capture the consensus of findings present in similar patient reports while allowing for patient-specific individuality and sparsity in the reporting of findings. We show through extensive studies across datasets, judge models and report generation models that such retrieval-based consensus-trained judge models improve the overall quality of automated reports by 8-17% as measured through comparison with ground truth reports on the reported findings in evaluation settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.