acceptodds
Under review as a conference paper at ICLR 2027

FUNCTIONAL COVERAGE FOR SELECTIVE MICROBIAL TRAIT PREDICTION UNDER TAXONOMIC SHIFT

Abstract

Functional-pathway similarity may help identify reliable predictions for microbial lineages absent from a model's labelled training set. We evaluate this hypothesis on 15,002 genome accessions using taxonomic holdouts, a fixed protein-embedding classifier, and three complementary coverage coordinates. In the initial family-held-out analysis, KEGG proximity ranks correct predictions above errors on three pathogenicity/biosafety targets with mean AUROC 0.652, compared with 0.531 for protein-embedding proximity. We then examine whether this diagnostic signal improves selective prediction, regenerating accession-level predictions and quantifying uncertainty across families. Combining pathway coverage with classifier confidence yields a mean balanced-accuracy advantage of 0.069 over confidence alone at 50% retention, but the paired family-bootstrap 95% interval [-0.055, 0.202] includes zero. Data-provenance checks further show that all negative pathogenicity labels are biosafety-derived proxies and that a small number of records have inconsistent or placeholder taxonomy. Across 468 refitted taxonomic-holdout models, aggregate phenotype AUROCs reproduce to within 0.00074. Our results distinguish informative coverage geometry from established selection utility: functional proximity is a useful exploratory diagnostic in this setting, but a robust reliability improvement is not demonstrated. We provide prediction-level supplementary results and reproducible statistical checks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.