PROTAX-SSL: Probabilistic Taxonomic Classification with Self-Supervised Representations
Abstract
Taxonomic classification of DNA barcodes is fundamental to biodiversity monitoring, yet remains challenging due to class imbalance, incomplete databases, and unknown taxa. In such settings, reliable taxonomic inference requires calibrated uncertainty estimates and identifying unknown taxa. Self-supervised foundation models have demonstrated high-order representations that capture taxonomic relationships, but downstream classifiers built on them typically lack principled uncertainty estimates and mechanisms for modeling unknown taxa. PROTAX, a probabilistic classifier from statistical ecology, provides calibrated predictions and explicitly models taxonomic structure complexities, but operates on raw sequences, limiting its accuracy relative to self-supervised approaches. We bridge these two lines of work with PROTAX-SSL, which couples self-supervised DNA barcode representations with probabilistic taxonomic inference. We evaluate variants of PROTAX-SSL across three datasets spanning diverse experimental settings. PROTAX-SSL improves calibration and unknown-species detection, while its modular design leverages complementary representations to improve accuracy. We provide a basis for reliable taxonomic analysis and potential species discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.