acceptodds
Under review as a conference paper at ICLR 2027

Do Genomic Language Models Inherit the Right Assumptions?

Abstract

Genomic language models largely inherit three choices from natural-language models: subword tokenization, learned input embeddings, and likelihood-based pretraining. We test whether these choices align with DNA using controlled comparisons that hold architecture, data, and nucleotide budget fixed, together with model-free diagnostics. Under DNABERT-2's published merge rules, BPE is mutation-stable: a substitution moves boundaries by a mean radius of bp, less than one mean token length. The larger mismatch is placement. At coding-sequence edges, boundaries occur at a length-matched null; the depletion replicates across eleven windows spanning six chromosomes and persists against a compositional control with the opposite predicted sign. Landmark stratification shows depletion at splice sites but enrichment at stop codons. Reverse-complement symmetry separates into segmentation, embedding, and contextual-representation levels. Five released embedding tables give ; only two have enough reverse-complement pairs for formal inference, and an architecturally equivariant model still has an unconstrained embedding table. Across four tokenizer arms, repeats receive – of total loss improvement versus – for the CDS-assigned class. In a controlled dose intervention, increasing repeat loss weight lowers retained splice-site signal in mean-pooled representations from to relative to a randomly initialised backbone across two seeds (blocked slope , CI ). Placebo arms make a batch-only explanation less consistent with the observed pattern. These results motivate evaluating inherited NLP design choices against genomic invariances and functional structure rather than assuming direct transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.