LPDP: Inference-Time Reward Control for Variable-Length DNA Generation with Edit Flows
Abstract
Reward-guided DNA generation is typically limited to fixed-length sequences, making insertions and deletions difficult to control. We introduce Local Perturbation Discrete Planning (LPDP), a training-free inference-time controller for variable-length Edit Flows with frozen reward oracles. At each guided step, LPDP exhaustively scores all valid one-edit actions, selects promising roots, and allocates a bounded lookahead budget to local continuations. The resulting future rewards are backed up to the root before a single edit is executed, enabling structured reward-guided search over variable-length sequence spaces without retraining either the generator or the reward model. Across enhancer and splice benchmarks, LPDP consistently improves predictive rewards over the evaluated controllers, with statistically significant paired improvements, while preserving reference-like sequence composition and motif statistics. Ablations reveal task-dependent mechanisms: root-action selection accounts for most of the gain on enhancer, whereas local lookahead is more important on splice. The improvements remain consistent across random seeds and independent re-scoring with ChromBPNet and MaxEntScan. Overall, LPDP provides an effective inference-time planning framework for reward-guided variable-length DNA generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.