DNA-RADIO: Towards an Agglomerative DNA Embedding Model
Abstract
DNA language models (LMs) achieve strong genomic prediction performance, yet their strengths remain distributed across models, with no single model consistently dominating. Through layer-wise linear probing and centered kernel alignment, we find that downstream performance often saturates early and deeper layers contribute little to additional downstream-relevant signal, suggesting that these complementary capabilities can be consolidated into a compact representation. We introduce DNA-RADIO, an agglomerative embedding model that distills heterogeneous DNA LMs into a single student up to 29.4 smaller than its teachers. We then enrich its representations through post-training on large-scale labeled sequences while preserving inherited knowledge of distillation. Under a frozen-feature linear evaluation protocol, distillation using only unlabeled sequences improves average performance by up to 15.85% over its teachers and other sequence-only DNA LMs on held-out sequences, establishing state-of-the-art results despite its substantially smaller size. Post-training yields a further improvement on BEND, a benchmark held out in its entirety with its sequences excluded from training. These results establish multi-model distillation and post-training as an effective route to compact, unified, general-purpose DNA representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.