Evolutionary Constraints as Data: Augmenting RNA Language Models with Phylogenetic Guidance
Abstract
Fine-tuning RNA language models is often limited by scarce labelled data. Data augmentation can help, but nucleotide edits may change the biological properties being predicted, making the original labels unreliable. We introduce PhyloAug, a plug-in framework that uses phylogenetic information to guide RNA sequence augmentation. PhyloAug uses relative evolutionary rates estimated from homologous sequences to select candidate editing sites, then uses a structure-conditioned RNA language model to generate replacement nucleotides. The resulting variants are used to fine-tune existing models without changing their architectures or training objectives. Across seven RNA functional prediction benchmarks and multiple model backbones, PhyloAug improves performance over unaugmented finetuning in most evaluated settings. Comparisons under matched augmentation and training budgets support the contribution of phylogenetically guided site selection. Complementary distributional diagnostics show smaller sequence-distribution shifts than less constrained editing policies. The benefits vary across tasks and label budgets, highlighting both the value and limits of evolutionary guidance for improving RNA language model adaptation with limited labelled data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.