DP-DNA: Dual-Policy Reinforcement Learning for DNA Discrete Diffusion Models
Abstract
Reinforcement learning (RL) provides a powerful framework for steering generative models toward task-specific objectives and has shown promising results in regulatory DNA design. However, fine-tuning discrete diffusion models with downstream rewards can lead to undesirable distributional degeneration, where improved predicted activity is accompanied by reduced pattern diversity, distorted sequence statistics, and decreased biological naturalness. In this work, we first systematically characterize this failure mode and find that its severity is strongly generation-order dependent: after RL, different unmasking schedules produce markedly different degrees of sequence collapse. Motivated by this observation, we formulate discrete diffusion generation as two coupled decisions: and . We introduce a dual-policy RL framework DP-DNA, which optimizes both a generation-order planner and a content generator with a joint order–content GSPO objective. A free-generation stream trains both policies to generate DNA with high functional reward, while a real-sequence stream teaches them to reproduce the sequence patterns of real DNA. Across HepG2, K562, and SK-N-SH enhancer design, our method achieves high predicted activity while producing more natural sequence distributions, supporting dual-policy learning for function-directed DNA generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.