acceptodds
Under review as a conference paper at ICLR 2027

MIRRA: Fast, High-Fidelity Molecular Generation with Chemistry-Informed Discrete Flow Matching

Abstract

Fast, high-fidelity molecular generation requires accurate chemical choices and efficient sampling. Chemical building blocks reduce graph size, but make each node a choice among tens of thousands of related identities. We introduce MIRRA (Molecular Identity Representation, Refresh and Assembly), a discrete flow-matching generator that learns from chemical knowledge and evaluates identity probabilities only when needed, while keeping every identity in its vocabulary available. Chemical features share information across related blocks, and learned identity corrections distinguish them in molecular context. By deciding which positions will refresh before computing their probabilities, the sampler avoids unused predictions while preserving each sampling transition. Against the strongest baselines in each task, MIRRA reduces chemical and graph distribution errors on ZINC250k by 29–31% and 53–60%, respectively, with generation up to 8.6x faster. Property targeting yields an estimated 7.8–20.8x more distinct novel hits per second. On a reserved polymer test, MIRRA achieves 100% periodic validity with known chain topology, versus 57.0% for a polymer variational autoencoder. These results demonstrate fast, high-fidelity generation with the full chemical vocabulary available throughout sampling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.