Benchmarking Symbolic Rule Discovery for Biological Motifs
Abstract
Existing methods for symbolic discovery primarily focus on recovering mathematical equations from numerical data, leaving it unclear whether they can induce valid logical rules from discrete sequences. We introduce SeqRule-Bench, a benchmark designed for symbolic rule discovery in biological motifs. Unlike conventional biological sequence analysis tasks formulated to classify or score sequences, it formulates each task as rule induction from labeled positive and negative sequences, requiring outputs in a human-readable domain-specific language. SeqRule-Bench integrates real biological data with synthetic tasks featuring ground-truth rules and informative negatives, addressing the scarcity of verified rules and the difficulty of obtaining negatives. To support systematic evaluation, we propose comprehensive metrics to analyze the trade-offs between accuracy, interpretability, and cost, including a novel boundary accuracy metric. By evaluating a broad range of methods, the benchmark reveals that even when these methods achieve high accuracy on some tasks, they still exhibit a gap in boundary accuracy and incur higher costs. The benchmark will be fully open-sourced upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.