acceptodds
Under review as a conference paper at ICLR 2027

WAVEBench: Assessing Generalization of Genome Language Models in Wastewater Virus Detection

Abstract

Genome language models (gLMs) are trained on DNA sequences to illuminate the relationship between genotype and phenotype. This goal is only useful if models are able to extend their reasoning beyond annotated data, or generalize out-of-distribution. However, this evaluation is confounded, since all biological sequences descend from the same origin. We introduce WAVEBench, a benchmark that withholds whole taxa, at ranks from species to class, from the labelled folds used to train a downstream probe. This way, each held-out taxon is novel to the probe. To test generalization, we devised a fold-assignment algorithm that, given any hierarchy of relatedness, produces a cross-validation partition where each fold's held-out data is stratified into progressively more distant novelty ranks from the remaining folds. We apply it to novel virus identification in wastewater metagenomics, a timely public-health surveillance task where the evolutionary distribution shift is intrinsic. We evaluated a logistic probe on 51 models from 11 gLM families and compared them to a 5-mer baseline. Median AUROC falls from species to class, but individual models do not decline uniformly; different ranks contain different clades, so these comparisons do not isolate an effect of evolutionary distance. Also, performance varies substantially relative to a k-mer composition baseline, with a subset of models achieving higher leaderboard AUROC. Overall, we contribute a fold-construction algorithm, a benchmark on a pressing public-health problem, and an assessment of gLMs' ability to generalize to taxa withheld from downstream training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.