Skip-RoPE: understanding long-range genomic interactions using skip-wise training
Abstract
Short-read sequencing generates genomic data at scale through parallel processing of millions of DNA fragments, typically spanning hundreds of base pairs, whereas underlying regulatory interactions can span millions. Genomic language models (gLMs) often bridge the gap by training on long contiguous inputs assembled from short reads, but this incurs significant data generation and computational cost. We propose a modification of Rotary Position Embeddings (RoPE), Skip-RoPE, that encodes genomic distances between packed fragments by using their aligned genomic coordinates to skip missing bases. This biologically-informed decoupling of context window size and genomic distance lets a 100M-parameter model trained to an effective context length of Mb using 60% fewer FLOPs than full-length training. Skip-RoPE also supports evaluation on discontinuous sequence inputs. We evaluate a 1 Mb model by predicting enhancer-gene interactions in a CRISPRi dataset through trained classifiers on frozen embeddings. Using Skip-RoPE, we achieve comparable results to full-context evaluation while requiring up to fewer input tokens. Apart from training and inference efficiency gains, we believe that this will enable a new class of gLMs that directly target raw short-read data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.