Do Genomic Language Models Inherit the Right Assumptions?
Abstract
Genomic language models largely inherit three choices from natural-language models: subword tokenization, learned input embeddings, and likelihood-based pretraining. We test whether these choices align with DNA using controlled comparisons that hold architecture, data, and nucleotide budget fixed, together with model-free diagnostics. Under DNABERT-2's published merge rules, BPE is mutation-stable: a substitution moves boundaries by a mean radius of bp, less than one mean token length. The larger mismatch is placement. At coding-sequence edges, boundaries occur at a length-matched null; the depletion replicates across eleven windows spanning six chromosomes and persists against a compositional control with the opposite predicted sign. Landmark stratification shows depletion at splice sites but enrichment at stop codons. Reverse-complement symmetry separates into segmentation, embedding, and contextual-representation levels. Five released embedding tables give ; only two have enough reverse-complement pairs for formal inference, and an architecturally equivariant model still has an unconstrained embedding table. Across four tokenizer arms, repeats receive – of total loss improvement versus – for the CDS-assigned class. In a controlled dose intervention, increasing repeat loss weight lowers retained splice-site signal in mean-pooled representations from to relative to a randomly initialised backbone across two seeds (blocked slope , CI ). Placebo arms make a batch-only explanation less consistent with the observed pattern. These results motivate evaluating inherited NLP design choices against genomic invariances and functional structure rather than assuming direct transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.