acceptodds
Under review as a conference paper at ICLR 2027

Self-Editing and Preference Distillation for Chest X-Ray Report Generation

Abstract

Reinforcement learning has advanced chest X-ray report generation, but assigning a single report-level reward can suppress correct findings in low-reward reports and reinforce local errors in high-reward ones. Self-refinement offers a way to address such local errors, yet existing approaches have limitations: inference-time methods can be unreliable without external feedback, while training-time methods rely on externally derived corrections or refinement supervision rather than explicitly optimizing the generation model itself to perform self-refinement. We propose ED-SR (Editing and Distillation for Self-Refinement), a training-time reinforcement learning framework in which a single vision-language model serves as both generator and editor. The generator produces candidate reports, while the editor re-examines the images and its generated drafts to identify potential errors, explain why they are problematic, and propose sentence-level revisions, deletions, and additions. ED-SR augments report-level policy optimization with two losses: (i) a counterfactual edit-reward loss that assigns fine-grained credit to self-generated editing actions using the reward change from applying each edit alone, and (ii) a correction-distillation preference loss that trains the generation policy to prefer reward-improving revisions over their original drafts. Because the generator and editor share parameters, edit-level supervision directly updates the same model used for report generation, while gradient conflict resolution manages competing report-level and preference updates. At inference time, ED-SR generates reports in a single pass without an additional refinement stage. Experiments demonstrate state-of-the-art performance across natural language generation and clinical efficacy metrics, achieving the best results on 9 of 11 metrics on MIMIC-CXR and all evaluated metrics on CheXpert Plus. Code will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.