acceptodds
Under review as a conference paper at ICLR 2027

TaxoSet: Open-Reference Taxonomic Assignment of DNA Barcodes

Abstract

A classifier can only return a name from its label space, and in DNA barcoding that space is the sequence reference, not the whole taxonomy. Our four curated corpora carry a barcode for of the species it names. For a query from any of the other , a method that can only return reference names has an accuracy of exactly zero, whatever the model. We introduce open-reference taxonomic assignment. The reference supplies the evidence, a query's nearest sequences and their lineages. The taxonomic hierarchy, 2.9 million taxa at all ranks, supplies the candidate names, sequenced or not. When the evidence locates a branch but cannot separate its unsequenced members, choosing one would invent information, so the method returns a set of names, whose coverage we never report without its size. TaxoSet implements this with a 0.34M-parameter network that learns how to decide, not what the taxa are, so it applies to names it never saw in training. Removing whole clades from the reference before retrieval makes novelty a label rather than an estimate. On ITS queries whose genus is absent, the returned set holds the true genus of the time at a median of 11 names. At species rank it falls to , where the limit is detecting novelty rather than the set itself. On a third-party benchmark whose novelty is real rather than simulated, TaxoSet reaches on ITS, where no published method returns the absent genus at all.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.