Learning the Language of the Microbiome with Transformers
Abstract
Self-supervised pretraining has become central to biological machine learning, yet microbiome data remains comparatively underexplored in terms of both modeling approaches and evaluation frameworks. To address this gap, we present Atlas, a pretraining dataset of 539,308 microbiome datapoints from the MGnify database. Using Atlas, we train the Waypoint family of microbiome foundation models: a series of GPT-2 style causal language models ranging from 6M to 170M parameters. We also introduce Compass, a curated benchmark of eight predictive tasks spanning biome classification, drug-microbiome interactions, drug degradation, and infant gut development. Using this benchmark, we compare the performance of Waypoint models against classical baselines and the existing Microbial General Model (MGM) foundation model. Our results show that pretraining consistently improves downstream task performance, that both dataset scale and tokenization strategy affect model quality, that pretraining supports favorable scaling behavior and that representations learned during pretraining transfer between microbiome domains. Furthermore, pretrained transformer models outperform classical methods on Compass tasks with roughly 10,000 or more training examples, a number that is attainable for modern microbiome studies. Finally, we demonstrate that the Waypoint models achieve state-of-the-art performance among publicly available microbiome foundation models. Overall, our work highlights the importance of large-scale self-supervised pretraining and tokenisation choices in this domain and establishes Atlas, Compass, and the Waypoint models as valuable resources for the research community in this emerging field.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.