MIP²: Information-Preserving Projector for Molecular Large Language Models
Abstract
Molecular large language models (LLMs) use cross-modal projectors to map pretrained molecular features into the language model embedding space. However, existing projectors do not explicitly enforce the recoverability of these features after projection. Therefore, the projectors may limit access to fine-grained features captured by molecular encoders. To address this challenge, we propose Molecular Information-Preserving Projector (MIP²), a molecular-language projector that combines feature preservation with flexible nonlinear adaptation. MIP² first integrates topological and geometric features through reversible coupling, preserving both encoder representations. It then separates feature preservation and nonlinear adaptation into orthogonal subspaces, yielding an injective projection with an explicit recovery map. MIP² further retains one token per atom and enables bidirectional attention within the molecular span to support atom-level structural interactions. Experiments on molecular question answering (MoleculeQA dataset), description generation and retrosynthesis (Mol-Instructions dataset), and property prediction (Mol-Instructions and eight MoleculeNet datasets) demonstrate the effectiveness of MIP². Backbone comparisons and ablation studies further support the benefits of the proposed design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.