OmniSpecGen: Where to Fuse Multimodal Spectra for Molecular Structure Prediction
Abstract
*De novo* molecular structure prediction from spectral measurements is a fundamental task in chemistry. Although existing models have primarily focused on single modalities or predefined modality sets, the measurements available in practice vary across molecules, calling for flexible use of complementary information from arbitrary subsets of modalities. The heterogeneous formats of these spectra motivate modality-specific processing, but the generation pipeline must transition to joint processing to generate a single molecule. This raises a central architectural question: *To best exploit complementarity, how much processing should remain modality-specific before joint processing begins?* To address this question, we introduce *OmniSpecGen*, a unified framework for comparing four fusion stages within a common diffusion architecture. Our comparison identifies fusion of modality-specific representations before pooling as the most effective strategy, with the resulting model outperforming the current state-of-the-art model for flexible spectrum-to-structure generation across the evaluated multimodal configurations. These findings clarify when to begin joint processing and provide an effective framework for molecular structure elucidation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.