Collapse, Repair, and Drift: Dissecting Fine Tuning Dynamics in Speech Language Models
Abstract
Supervised fine tuning (SFT) is commonly monitored through training and evaluation loss, but loss does not fully characterize how model functionality evolves along the same trajectory. We study this gap through dialect adaptation of Qwen3-ASR, MiMo-Audio, and FireRed-LLM across learning rates from 10⁻⁵ to 10⁻³. Character error rate (CER) reveals pronounced transient collapse and recovery whose severity and temporal structure are only weakly reflected by loss. Bidirectional layer patching reveals strong cross checkpoint functional incompatibility during recoverable collapse, with model dependent depth profiles that reorganize during recovery. Pairwise gradient similarity alone does not reliably indicate optimization stability, whereas raw gradient spectral analysis provides a clearer structural distinction between stable and unstable optimization: successfully converged runs approach a broad and temporally stable spectral energy distribution, while failed runs remain spectrally concentrated or exhibit persistent fluctuations. Finally, AdamW update transfer bypasses pronounced early collapse, yet substantial Mandarin degradation still emerges later in training. These results motivate three separable SFT dynamics: collapse, repair, and drift, and show that transient collapse is not necessary for persistent source domain degradation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.