Quality Before Quantity: What Governs the Gain from Self-Training in Low-Resource Speech Recognition
Abstract
Self-training with pseudo-labels is a commonly used strategy for domain adaptation, yet practitioners have no evidence with which to predict how much benefit it will provide, or how aggressively to filter, before running it. We study confidencefiltered self-training on speaker-disjoint and fully cross-corpus target splits, using an external test corpus that is never pseudo-labelled, oracle ceilings obtained from true target labels, and ten seeds per condition with paired non-parametric tests. We perform the full threshold sweep twice, using two different pools of unlabelled spontaneous Hindi 6.5 hours sampled speaker-disjointly from the target corpus, and 30 hours from a different corpus altogether and then perform a second self-training round on top of them. For both pools, self-training improves word error rate by 6.5 points absolute (15% relative), improves an external benchmark it never saw by 2.3 and 3.1 points, does not harm source-domain accuracy, and yields a student that outperforms its own teacher, trained on 128 hours of labelled speech, on every seed. Sweeping the threshold over the two pools reveals the same structure in both: once retained volume is controlled, downstream error is perfectly rank-ordered by the error rate of the retained pseudo-labels, and the optimum is interior, falling at the same threshold in both pools even though one is five times the size of the other and the teachers differ by fifteen points in quality. The relationship is strongly concave: almost all of the benefit comes from removing the worst labels, and filtering severely enough to retain only a small fraction of the data reverses the benefit even though those are the cleanest labels in the study. The second self-training round saturates: the improved teacher creates labels 2.8 points cleaner overall but only 0.4 points cleaner at the operating threshold, because iteration repairs precisely the labels that filtering had already discarded. We argue that filtering and iteration are substitutes rather than complements, that the pre-registered slope transfers in ordering but not in magnitude across corpora, and that quality drives the result, with diminishing returns, until volume becomes the limiting factor.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.