Prefix-Conditioned Boltzmann Generators: Towards General Molecular Sampling
Abstract
Reliable estimates of molecular properties require sampling the equilibrium conformational ensemble, which remains a bottleneck for molecular simulations. Boltzmann Generators (BGs) address this by utilizing generative models with tractable likelihoods to enable importance sampling towards the Boltzmann distribution. Recently, efforts have been made to scale BG training to increasingly large and diverse data. However, current BG approaches are often optimized for one particular chemistry, and only individual molecules. We introduce Prefix-Conditioned Boltzmann Generators (PBGs), an autoregressive Boltzmann generator that conditions directly on text and coordinate prefixes rather than a specialized atom-wise molecular encoder. Across a series of models trained on Many Peptides, we identify scaling relationships linking sampling quality to model size, training tokens, training compute, atom count, and coordinate discretization resolution. Guided by these relationships, we scale PBGs up to 640M parameters, achieving state-of-the-art performance on unseen peptide systems up to 157 atoms. We observe that models trained across diverse molecular classes and environmental conditions, achieve strong performance in our main benchmark at comparably lower training cost. This performance suggests shared information and potential transferability across molecular types. Beyond peptides, PBGs' prefix conditioning enables flexible generation across molecular systems and environments, including small-molecule sampling, solvation free energy estimation, and conditional microsolvation of systems with over 280 atoms. Together, these results establish PBGs as a promising route toward general molecular sampling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.