Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
Abstract
Automatic speech recognition degrades nonlinearly in the wild: as acoustic corruption compounds, local confusions give way to omissions, repetitions, empty outputs, and semantically plausible but unsupported transcriptions. We present MEGA-ASR, a learnability-aligned framework that treats this transition as a linked problem of coverage, grounding, and optimization. First, VOICES-IN-THE-WILD-2M composes seven atomic acoustic effects into 54 controllable scenarios, providing hard but learnable supervision for compound environments. Second, Acoustic-to-Semantic Progressive Supervised Fine-Tuning (A2S-SFT) establishes acoustic grounding before adapting semantic recovery: it first trains the acoustic interface, then the language-side decoder, and finally aligns both end to end. Third, Dual-Granularity WER-Gated Policy Optimization (DG-WGPO) starts from this grounded policy and changes its preference with the error regime, favoring token-level refinement when speech remains recoverable and structural reconstruction when it does not. The three stages therefore form a capability chain rather than independent training add-ons. Compared to the baseline, Qwen3-ASR-1.7B, MEGA-ASR achieves 45.69% versus 54.01% WER on VOiCES R4-B-F and 21.10% versus 29.34% WER on NOIZEUS Station-0dB. On the mixed real/simulated compositional settings, it reduces WER by 65.8%/52.5% relative to Gemini-3-Flash and by 69.4%/69.1% relative to Whisper Large-v3.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.