Geometry Is Not Enough: Ranking RNA Structure Candidates Against a Sequence Prior
Abstract
Computational RNA structure prediction produces many candidate conformations per sequence, and selecting among them remains a bottleneck, especially for long RNAs, whose candidates often share correct base-paired stems but differ in how those stems pack in three dimensions. Geometry alone can confirm that a packing is physically plausible, but not that it is the one the sequence adopts. We therefore frame evaluation as sequence–geometry alignment and introduce SIRGE, an SE(3)-invariant evaluator that conditions the geometry of every nucleotide pair on attention maps from a frozen RNA language model, used as a sequence prior. A contact-recovery objective trains these pair representations to recognise native tertiary contacts, so SIRGE judges each candidate's packing against what its sequence implies. Against six established evaluators under two distribution shifts, SIRGE ranks candidates best. On CLS-1, a benchmark of held-out, longer RNAs with broader candidate quality, it improves Kendall- alignment by 70% over the strongest baseline; on candidates from six unseen deep-learning predictors, it almost doubles the strongest baseline's top-1 hit rate. Ablations show that the prior and the contact objective are complementary, and that alignment helps most on long RNAs. An interpretability study shows that the prior carries tertiary-contact information beyond base pairing, and that SIRGE's scores track native tertiary packing where a geometry-only counterpart does not. These results suggest that for long RNAs geometric plausibility is necessary but not sufficient, and that pretrained sequence models can supply the missing reference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.