TaxoSet: Open-Reference Taxonomic Assignment of DNA Barcodes
Abstract
A classifier can only return a name from its label space, and in DNA barcoding that space is the sequence reference, not the whole taxonomy. Our four curated corpora carry a barcode for of the species it names. For a query from any of the other , a method that can only return reference names has an accuracy of exactly zero, whatever the model. We introduce open-reference taxonomic assignment. The reference supplies the evidence, a query's nearest sequences and their lineages. The taxonomic hierarchy, 2.9 million taxa at all ranks, supplies the candidate names, sequenced or not. When the evidence locates a branch but cannot separate its unsequenced members, choosing one would invent information, so the method returns a set of names, whose coverage we never report without its size. TaxoSet implements this with a 0.34M-parameter network that learns how to decide, not what the taxa are, so it applies to names it never saw in training. Removing whole clades from the reference before retrieval makes novelty a label rather than an estimate. On ITS queries whose genus is absent, the returned set holds the true genus of the time at a median of 11 names. At species rank it falls to , where the limit is detecting novelty rather than the set itself. On a third-party benchmark whose novelty is real rather than simulated, TaxoSet reaches on ITS, where no published method returns the absent genus at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.