Actionable Feedback: Improving Model Capabilities From User Edit Data on Clinical Documentation
Abstract
Large language models (LLMs) are increasingly being deployed in clinical documentation workflows, where clinicians review and edit their outputs before use. In principle, these clinician edits represent a scalable source of supervision for model training. In practice, not all edits represent actionable feedback: for instance, on data drawn from real-world deployment of an AI clinical documentation product, our automated review suggests that over 45% of edited outputs contain user-added information that is not directly supported by the inputs to the model. Naively fine-tuning on such edits risks teaching the model to hallucinate. In this paper, we develop a method to address this challenge. Given that real clinician edits on deployed documentation tools are not publicly available, we first develop a taxonomy of clinician-editing behaviors, informed by data from a deployed AI clinical documentation product. Then, we develop a method for generating semi-synthetic clinician-edit data, informed by our taxonomy, on a discharge-summary generation task based on MIMIC-III, a public dataset. Finally, we introduce simple filtering methods for identifying actionable edits, with accuracy evaluated by expert review. We show empirically, on our semi-synthetic dataset, that naive fine-tuning on raw edited output increases ungrounded content, while our approach of progressively filtering toward grounded edits enables improvement on clinician-alignment metrics while retaining faithfulness to the input context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.