acceptodds
Under review as a conference paper at ICLR 2027

Beyond Single Variants: A Benchmark for Regulatory Epistasis in Genomic Sequence Models

Abstract

Genomic sequence models, both supervised sequence-to-function predictors (Enformer, Borzoi, AlphaGenome) and self-supervised genome language models (Nucleotide Transformer, Evo 2), are increasingly used to interpret non-coding variation, yet they are benchmarked almost exclusively on variants tested one at a time against a reference background. Real haplotypes carry many variants per regulatory element and regulatory grammar is non-additive, so the practical question is whether models can compose variant effects. We assemble a benchmark from MPRA diplotype experiments that assay all four haplotypes (REF, A, B, AB) of thousands of variant pairs, together with  1.15M single-variant (variant × cell-type) measurements across three studies, and evaluate five models along a hierarchy of tasks: single-variant regression (T1), interaction regression (T2), non-additive pair detection (T3), and full-versus-additive prediction across 421 phased personal genomes (T4), each under a zero-shot and a frozen-encoder linear-probe readout. On Siraj et al. (2026), models rank confidently measured single-variant effects moderately well (zero-shot Spearman ρ up to 0.42; linear probe up to 0.73) but capture far less of the measured interaction signal (ρ ≤ 0.32; AUROC ≤ 0.68, against 0.62 for inter-variant distance alone). Linear probes recover substantial single-variant signals from gLM representations that their likelihoods do not expose, yet bring interaction detection at most to the level of the distance baseline. On personal genomes, full-haplotype predictions are no more accurate than additive reconstructions from each model's own marginal effects. Current models' interaction predictions thus add little beyond their marginal effects, a gap that matters for variant prioritization on real haplotypes and for sequence design, and that motivates interaction-aware training data, objectives and benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.