acceptodds
Under review as a conference paper at ICLR 2027

Evolutionary Informed Translation from Protein to Nucleic Acid Sequence

Abstract

The choice of nucleic acid sequence for a given protein strongly impacts function and is species-specific. Current protein-to-nucleic-acid translation models treat sequences in isolation, often over-favoring frequently used codons. To generate more natural nucleic acid sequences, we introduce PINATA (Phylogenetically Informed protein to Nucleic Acid Translation with sequence Alignment), a model that leverages sequences from related species encoded in multiple sequence alignments (MSAs) to guide prediction. Our architecture leverages evolutionary context and learns species-and protein specific context to make biologically-informed predictions. Evaluated across 61,934 gene trees and 882 species, PINATA generates sequences that mirror the naturally-occurring sequence substantially better than existing language models. PINATA leverages MSA data to improve exact native codon recovery by 17.95 percentage points over existing language models, effectively preserves natural rare-codon nuances, and more closely matches native sequence properties. PINATA maintains strong predictive performance even with limited evolutionary data. While previous methods largely treated sequences from different species in isolation, our method leverages the evolutionary relationships between nucleic acids and proteins across species to create a more biologically-informed translation. We demonstrate that leveraging rich evolutionary data through an biologically-designed method strongly aids in translating protein sequences to nucleic acid sequences across species. We will make data and code publicly available upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.