CAMS-Diff: Coupled Adaptive Mixed-Space Diffusion for Molecular Translation
Abstract
Diffusion models are increasingly used for molecular translation, including molecule captioning, text-guided generation, forward reaction prediction, and retrosynthesis. Continuous sequence diffusion provides a common formulation across these tasks, with molecules represented through SMILES. However, two limitations remain. (1) At low noise levels, continuous representations can retain substantial local token information, reducing reliance on broader sequence context and the input condition. (2) During inference, uniform timestep reduction can also skip informative regions of a non-uniform denoising trajectory. We introduce Coupled Adaptive Mixed-Space Diffusion (CAMS-Diff). During training, CAMS-Diff couples token-wise adaptive Gaussian noising with a discrete representation-masking process governed by the same learned schedule, selectively masking direct local token evidence. During inference, it profiles reconstruction changes in token-wise log-SNR space and uses the resulting trajectory importance to select denoising timesteps under a fixed sampling budget. Across five molecular translation and generalization settings, CAMS-Diff improves molecular and semantic recovery; on ChEBI-20, it reaches 0.916 MACCS compared with 0.883 for BiMol-Diff. Controlled ablations isolate both mechanisms: coupling outperforms adaptive noising, masking, and their uncoupled combination, while trajectory-informed timestep selection outperforms uniform subsampling at matched budgets. Two-step captioning, using two denoiser evaluations instead of 2000, achieves quality comparable to full-step BiMol-Diff with an approximately 245x wall-clock speedup. CAMS-Diff also transfers zero-shot to PCDes, improving MACCS from 0.763 to 0.810.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.