acceptodds
Under review as a conference paper at ICLR 2027

Ariadne: Regulatory DNA Design with Adaptive Reward Shaping and Reinforcement Learning

Abstract

Cis-regulatory elements (CREs), including enhancers, regulate gene expression through the composition and organization of transcription factor binding sites (TFBSs) within their sequence context. The vast sequence space makes exhaustive experimental screening infeasible, motivating computational methods to design enhancers with high activity in a target cell type. A central challenge is to improve predicted activity while determining whether optimization narrows the generated sequence distribution and whether gains extend beyond the predictor used for training. We present Ariadne, a reinforcement-learning framework that combines an enhancer sequence prior with target-cell-directed optimization to address this challenge. Ariadne first adapts an autoregressive DNA language model to massively parallel reporter assay (MPRA)-derived enhancer sequences through a single shared supervised fine-tuning (SFT) stage and then trains separate target-cell policies with proximal policy optimization (PPO), using the corresponding Enformer prediction as the reward signal. To account for changing reward distributions during optimization, the evaluated reward combines exponential moving average (EMA) normalization, within-batch ranking, and a clipped bonus above an EMA-tracked upper-quantile threshold. Normalization controls favor an evolving reward reference over fixed SFT statistics, but do not establish a universal advantage of EMA over batch normalization. Component ablations support modest incremental Enformer gains from rank, whereas tail has cell-dependent returns that are not uniformly reproduced by Malinois. PPO improves the reported Enformer and Malinois summaries relative to the shared SFT policy, but Hamming diversity decreases from SFT and Enformer-embedding concentration increases. These results support Ariadne as a computational pipeline for improving predicted enhancer activity and characterizing the accompanying distributional trade-offs. They do not establish universal superiority of the complete EMA–rank–tail reward, experimental activity, or cell-type selectivity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.