acceptodds
Under review as a conference paper at ICLR 2027

BasePrompt: Context as Prompt for Causal Genome Language Models in Zero-Shot RNA Fitness Prediction

Abstract

Causal genome language models (GLMs) are pretrained on long genomic context, whereas zero-shot RNA fitness benchmarks often expose isolated mature transcripts with little usable local context. We study whether a frozen causal GLM can use synthetic context to improve prediction without task-specific training. BasePrompt is a training-free procedure that constructs short synthetic 5′ and 3′ flanks around a target sequence. Its key mechanism is reverse-complement (RC) decoding, which opens a left-context channel for an otherwise forward-only model: we decode on the RC strand to synthesize an upstream flank, map back to the original strand, and then decode a downstream flank autoregressively. The main evaluation uses a shared fixed budget of three bases per side and reports matched RC-ensemble and direct-scoring results separately. On RNAGym, this fixed-budget comparison improves all three Evo2 backbones in the main text; for example, Evo2 7B rises from 0.276 to 0.292 Spearman ρ under matched RC scoring. Gains persist without RC ensembling. Assay-level and type-level decompositions are used as descriptive robustness views, making the type-averaged benchmark aggregation and the remaining heterogeneity explicit. Auxiliary DNA controls show that extension content matters in the tested settings, without establishing the causal mechanism of the RNA gains. We present BasePrompt as an inference-time context-construction method for causal genome language models, not as a reconstruction of true biological flanks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.