VeriSpeech-Repair: Geometry-constrained local repair with clean-context audits
Abstract
A speech editor may recover a missing word while changing audio outside the intended edit. We study this tension with VeriSpeech-Repair: a fixed set of local speech patches, an explicit no-repair choice, and separate tests for transcript recovery and preservation outside the original corruption span. On controlled LibriSpeech corruptions, broad patch selection roughly halves automatic word error rate (WER) but reduces exact outside-span context by an absolute 22.1% among outputs that can be aligned to the original audio. A score-free geometry rule preserves exact context under local splicing and lowers WER by an absolute 2.8% on 782 sources after excluding earlier calibration sources. In a separate comparison on those sources, the configured systems differ in architecture, input information, and proposal count. dots.tts.edit achieves 52.259% automatic transcript pass on corrupted cases, compared with 28.261% for the risk controller. Only five of the editor’s 2,346 corrupted outputs align to the original time map; none has exact context, and the other 2,341 cannot be assessed by this fixed-map measure. Transcript recovery and fixed-map preservation therefore give different outcomes for these systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.