acceptodds
Under review as a conference paper at ICLR 2027

Beyond Tokens: Read and Predict N-gram Spans with a Shared Dictionary for Genomic Foundation Models

Abstract

Large language models lock in a fixed tokenizer and next token prediction before pretraining, fixing the granularity, boundaries, and horizon of what they read and predict. This is restrictive for DNA, in which the relevant signals are sparse, overlapping, multiscale, and context-dependent. To exploit the compact alphabet and the informative chunks intrinsic to DNA, we introduce Beyond Tokens, a coadaptive dictionary of n-grams that enables autoregressive models to read and predict high order patterns of variable length and arbitrary position on top of a base level interface. Rather than redefining the sequence units, we treat each base token as an anchor for two directional operations: contextual pattern reading and future pattern prediction. On the input side, Spanreader retrieves context-relevant n-grams of variable length and offset via sparse Top K routing and adds the routed residual at the backbone input. On the output side, Spanpredictor forecasts multihorizon n-grams as atomic targets through a length-biased confidence gate operating over multiscale chunk losses, providing higher order supervision. The two modules share the same dictionary parameters, allowing the patterns employed for reading and those employed for prediction to coevolve throughout pretraining. We instantiate Beyond Tokens within a 1.2B genomic model and evaluate it through supervised frozen-representation probing, where it improves the matched-budget baseline and attains the highest mean among the compared models, surpassing the reported means of 10B and 40B baselines trained on more tokens. These results indicate that models can adaptively learn higher order tokens beyond a fixed vocabulary, jointly for reading and prediction, a capability naturally suited to small-alphabet domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.