acceptodds
Under review as a conference paper at ICLR 2027

Cumulative Acoustic Variation for Training-Free ASR Data Pruning

Abstract

Training-data pruning for automatic speech recognition (ASR) aims to reduce training-data requirements, yet model-dependent selection criteria can incur sub- stantial computational overhead. To address this, we propose a model-free acous- tic pruning approach that ranks training utterances using cumulative statistics computed directly from speech signals. Specifically, we develop three utterance- level criteria based on accumulated spectral entropy, spectral flux, and log-Mel temporal variation, requiring neither pruning-specific model training nor learned- model inference. Under a fixed utterance-count budget, these criteria retain both the extent and the local acoustic characteristics of each utterance through cumula- tive scoring. Experiments with Conformer on LibriSpeech show that the proposed acoustic criteria achieve the best performance at 5% and 1% retention and re- main competitive with the best-performing Teacher–Student criterion at 30% and 10% retention, while reducing data-selection time from 13.73 hours to less than 0.5 hours without requiring GPUs. In particular, Log-Mel Temporal Variation achieves the lowest WER at both 5% and 1% retention; at 5%, it reduces WER relative to random selection by 26.1% and 15.3% on test-clean and test-other, re- spectively. These results demonstrate that simple cumulative acoustic statistics provide an efficient and effective alternative to model-dependent criteria for ASR training-data pruning, with especially strong performance under highly restrictive retention budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.