LEARNING FROM USER EDITS IN DICTATION SYSTEMS: ATTRIBUTION AND ADAPTATION
Abstract
Users of deployed multi-stage systems edit what the system produced, and it is tempting to learn from every edit as if it were a correction. We test this assumption on a production dictation pipeline (speech recognition followed by LLM post-processing) with 900 edits from 45 users. In a non-random human-labeled subset of 359 edits, 47.9% added or removed content rather than correcting a mishearing. The cause of an edit can be estimated: an LLM that sees the logs and an independent transcript of the recording assigns the same category as a human annotator for 81.0% of held-out edits, against 50.2% for always predicting the most common category, although most labels were assigned with a transcript-based suggestion visible. In an adaptation study scored against proxy references, vocabulary learned from edits reduced errors on corrected utterances by 2.2–3.0 points, whereas negative examples learned from every edit increased errors, an increase a small hand-transcribed subset did not confirm. Routing updates by the estimated cause never beat learning vocabulary from every edit. Updates learned from edits need their own validation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.