Garak-2: Selective Low-Rank Compression of a Korean Speech Foundation Model
Abstract
I introduce Garak-2, an efficiency-oriented successor to Garak-1, a Korean speech recognition foundation model trained on multi-domain Korean speech. Garak-1 initializes its cache-aware FastConformer encoder from the March 2026 NVIDIA Nemotron Speech Streaming 0.6B checkpoint and replaces the original RNN-T decoder with a Token-and-Duration Transducer (TDT) decoder. Both the encoder and decoder are trained on Korean speech, with the decoder initialized from scratch. Starting from this Korean-trained Garak-1 checkpoint, Garak-2 targets the encoder, which dominates model parameters and CPU inference time. I apply rank-512 factorization to the feed-forward projection matrices in the middle 12 layers of the 24-layer encoder, preserving the original expansion width, nonlinearities, cache-aware attention and convolution modules, and TDT formulation. This reduces the parameter count from 622.0 million to 546.5 million, a 12.1% reduction (546.8 million with the small correction path of the evaluated model). To account for the effects of additional training, I compare the compressed model with an uncompressed Garak-1 baseline under matched continued-training budgets. Preliminary experiments across 19 Korean speech datasets show that a short period of continued training brings the compressed model's macro-averaged character error rate close to that of the equally trained baseline, with a small remaining deficit at matched training steps (+0.04 CER points) that varies with the checkpoint. The same analysis reveals dormant feed-forward modules in the middle layers: replacing 12 of them by weight-derived constants removes 16.2% of the parameters without any training, at a small accuracy cost (about 0.5% relative on a held-out set). CPU inference benchmarking at batch size eight with FP32 computation on eight cores shows a 16–18% reduction in total processing time. These initial findings suggest that selective encoder compression is a promising approach to reducing the size and inference cost of Garak-1 at a small cost in average recognition performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.