acceptodds
Under review as a conference paper at ICLR 2027

Steering Flexible-Size Molecular Generation via RL on Trans-Dimensional Flows

Abstract

Flexible-size generative models for structure-based molecular design allow size to adapt to the target objective, opening opportunities to out-of-distribution scientific discovery. However, steering such trans-dimensional processes with reinforcement learning (RL) requires accounting for actions that can change the state's dimension along the generative trajectory, such as insertions and deletions. We introduce Transdimensional Flow Policy Optimization (TFPO), formulating RL finetuning of flexible-size generative models as a parameterized-action Markov decision process, in which discrete structural decisions select an action-space component and where variables such as continuous positions and discrete atom types and bonds are sampled conditionally within it. We instantiate TFPO on Morph, a trans-dimensional 3D molecular generator and derive its corresponding per-step policy density. Across three settings, reward optimization steers both molecular properties and the model’s structural decisions. On QM9, we steer the model far beyond the pretraining size support, generating viable structures larger than the training support mean. On GEOM-Drugs, RL with verifiable rewards increases strict validity from 92.9% to 98.0% while decreasing the strain energy. For pocket-conditioned ligand generation, fine-tuning improves binding efficiency and successfully adapts ligand size to the target protein pocket.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.