acceptodds
Under review as a conference paper at ICLR 2027

Reconciling Coarse Tokenization with Single-Nucleotide Resolution in Genomic Language Modeling

Abstract

Genomic language models face a fundamental trade-off: coarse tokenization enables efficient long-context modeling, yet standard token-level objectives obscure single-nucleotide variation. We introduce Factorized Nucleotide Supervision (FNS), which derives position-specific nucleotide likelihoods from a distribution over -mers through exact marginalization. FNS retains the computational efficiency of non-overlapping -mers while providing dense nucleotide-level supervision and directly supporting variant scoring. We instantiate FNS in GENERator-v2, a family of autoregressive models for eukaryotic and prokaryotic genomes. Across training-free and supervised benchmarks, GENERator-v2 consistently improves over its predecessor, matches or exceeds substantially larger baselines, and offers up to faster inference. Beyond these model-level gains, our results challenge two common assumptions in GFM evaluation: supporting a longer context does not guarantee that a model benefits from the added sequence, and single-nucleotide tokenization is not required to learn nucleotide-scale biological structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.