Reconciling Coarse Tokenization with Single-Nucleotide Resolution in Genomic Language Modeling
Abstract
Genomic language models face a fundamental trade-off: coarse tokenization enables efficient long-context modeling, yet standard token-level objectives obscure single-nucleotide variation. We introduce Factorized Nucleotide Supervision (FNS), which derives position-specific nucleotide likelihoods from a distribution over -mers through exact marginalization. FNS retains the computational efficiency of non-overlapping -mers while providing dense nucleotide-level supervision and directly supporting variant scoring. We instantiate FNS in GENERator-v2, a family of autoregressive models for eukaryotic and prokaryotic genomes. Across training-free and supervised benchmarks, GENERator-v2 consistently improves over its predecessor, matches or exceeds substantially larger baselines, and offers up to faster inference. Beyond these model-level gains, our results challenge two common assumptions in GFM evaluation: supporting a longer context does not guarantee that a model benefits from the added sequence, and single-nucleotide tokenization is not required to learn nucleotide-scale biological structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.