acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Foundation Models for Microbiome Phenotype Prediction: What Actually Matters?

Abstract

Foundation models for the microbiome increasingly rely on large-scale pretraining and Transformer encoders, but it remains unclear whether these choices address the main difficulties of phenotype prediction from sparse abundance profiles. We revisit this assumption through a controlled study of inflammatory bowel disease (IBD) prediction that disentangles pretraining, architectural inductive bias, representation complexity, and cohort shift. A pretrained 34M-parameter MGM encoder and a 1.1M-parameter embedding model without any pretraining reach comparable held-out accuracy (92.8% versus 92.3%), indicating that neither scale nor pretraining alone determines downstream utility. Permutation-invariant models are similarly strong: DeepSets and a minimal learned-embedding model reach 91.9% and 92.3% in the same structural comparison. Under leakage-robust grouped cross-validation, however, classical models remain highly competitive: HistGradientBoosting reaches 0.979 AUROC, compared with 0.957 for the learned embedding, and a simple presence-vector MLP reaches 0.969. The decisive difficulty appears under cohort shift. When models trained on one cohort are evaluated on independent cohorts, performance drops sharply; on the balanced external cohort, AUROC ranges from 0.621 for the embedding model to 0.668 for Random Forest, far below the 0.974 in-family result. Cohort-probe analyses show that study identity is highly decodable from every representation (cohort accuracy 0.93–0.97 versus disease accuracy 0.76–0.82), and remains so after balancing disease composition, indicating that learned representations are dominated by cohort-associated structure. Crucially, adding a single additional cohort to the training set raises disease identification on an independent cohort from 32% to 49%, whereas gradient-reversal debiasing and model scaling do not. These results indicate that cross-cohort generalization in microbiome phenotype prediction is governed not by model capacity, pretraining, or explicit debiasing, but by the cohort coverage of the training distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.