Prediction Targets and Vocabulary Coverage in Bacterial Pangenome Models
Abstract
Genome language models over bacterial pangenomes predict each gene as one of a fixed set of protein families; we compare nine prediction targets for such models under a matched corpus, backbone, masking scheme, token budget and evaluation. Contrary to our pre-registered primary hypothesis, a head tied to a frozen protein encoder's centroids has higher held-out cross-entropy than a free classifier, by 0.225 nats. The gap depends on distribution shift. It is 0.435 nats for unseen species of a seen genus and indistinguishable from zero at unseen phylum, where resampling phyla rather than species widens the bound to 0.051 nats and the species-level reversal seen with a 302M-parameter backbone does not survive. Part of the gap reflects head size: at matched trainable parameters, the tied head is ahead. The target further sets what a model can represent. Out-of-vocabulary genes are 21.25% of held-out gene mass, and a closed-vocabulary head cannot emit their labels at all; a tied head can, but its exact-family accuracy on them is near zero and its net advantage depends on how finely the vocabulary is extended. An auxiliary target predicting cross-taxon gene persistence works on either head at small cost. Finally, frozen-probe conclusions under the widely used protocol depend on the layer probed: at the final layer no target we tested beats a protein-only baseline, while at the first block every arm's mean does on essential genes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.