acceptodds
Under review as a conference paper at ICLR 2027

PlantBGC: Label-Scarce Cross-Domain Representation Learning for Plant Biosynthetic Gene Cluster Discovery

Abstract

Plant biosynthetic gene cluster (BGC) discovery poses a label-scarce cross-domain learning problem: abundant microbial supervision is available, whereas plant BGC annotations are limited. We present PlantBGC, which combines a cross-domain representation learning framework with a standardized benchmark and evaluation protocol for plant BGC discovery. PlantBGC learns BGC-relevant Pfam context from microbial, adapts the representation to unlabeled plant genomes through masked language modeling, and further calibrates predictions through weak supervision derived from GO/KEGG annotations. On 34 curated plant BGCs, self-supervised plant adaptation increases strict 100% recovery from 29.4% to 67.6%; at matched genomic coverage, Stage 2 retains 61.8% recovery at 0.53%, versus 29.4% at 0.52% for Stage 1. Weak calibration further reduces residual primary-metabolism-like proxy burden beyond hard rule-only filtering without reducing curated-BGC recovery. Under matched per-genome candidate budgets, PlantBGC retains a 4.0-fold higher strict recovery than plantiSMASH while producing more compact candidate intervals. PlantBGC provides both an effective transfer framework and a controlled benchmark for label-scarce plant BGC discovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.