Adapting Language Models for Microbiome Prediction and Scientific Reasoning
Abstract
The human gut microbiome plays a vital role in shaping metabolism and immunity. To support research in this domain, AI systems must be able to reason about general microbiome knowledge and provide analysis of individual microbiome measurements, for example predicting disease status from a sample. Current models are suboptimal: microbiome-specific models are inexpensive and can classify individual samples, but are not designed for open-ended knowledge tasks; frontier language models have scientific knowledge and few-shot classification abilities, but are expensive to run. In our work, we address this discrepancy by curating data to help inexpensive pretrained language models support both of these capabilities. Starting with sample classification, we introduce a continued-pretraining (CPT) dataset that combines information from sample metadata and scientific literature. We show that while current language models can be adapted for sample classification, training on our dataset improves performance. To evaluate microbiome knowledge, we introduce an expert-authored benchmark of challenging questions that require critical assessment of scientific situations. Finally, to support both sample classification and scientific knowledge, we curate a multi-task finetuning data mixture using data derived from biological databases, publications, and sample-level classification. A 1.7B model trained on this mixture outperforms few-shot frontier models by 7 percentage points on average on our sample classification tasks and outperforms a 31B model by 10 points on one of our two knowledge tasks. Our results illustrate the feasibility of curating data that supports predicting biologically meaningful sample properties and scientific knowledge for a complex scientific domain within a single, compact model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.