Beyond Tokens: Read and Predict N-gram Spans with a Shared Dictionary for Genomic Foundation Models
Abstract
Large language models lock in a fixed tokenizer and next token prediction before pretraining, fixing the granularity, boundaries, and horizon of what they read and predict. This is restrictive for DNA, in which the relevant signals are sparse, overlapping, multiscale, and context-dependent. To exploit the compact alphabet and the informative chunks intrinsic to DNA, we introduce Beyond Tokens, a coadaptive dictionary of n-grams that enables autoregressive models to read and predict high order patterns of variable length and arbitrary position on top of a base level interface. Rather than redefining the sequence units, we treat each base token as an anchor for two directional operations: contextual pattern reading and future pattern prediction. On the input side, Spanreader retrieves context-relevant n-grams of variable length and offset via sparse Top K routing and adds the routed residual at the backbone input. On the output side, Spanpredictor forecasts multihorizon n-grams as atomic targets through a length-biased confidence gate operating over multiscale chunk losses, providing higher order supervision. The two modules share the same dictionary parameters, allowing the patterns employed for reading and those employed for prediction to coevolve throughout pretraining. We instantiate Beyond Tokens within a 1.2B genomic model and evaluate it through supervised frozen-representation probing, where it improves the matched-budget baseline and attains the highest mean among the compared models, surpassing the reported means of 10B and 40B baselines trained on more tokens. These results indicate that models can adaptively learn higher order tokens beyond a fixed vocabulary, jointly for reading and prediction, a capability naturally suited to small-alphabet domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.