When Does the Auxiliary Encoder Help in Multimodal Molecular Pretraining?
Abstract
Multimodal molecular pretraining often improves graph representations by aligning them with a second view of the same molecule, such as a SMILES string or 3D conformation. These gains are usually attributed to information learned by the auxiliary encoder. We show that this explanation is often incomplete: in several settings, a frozen, randomly initialized auxiliary network recovers much of the benefit of a trained one. We test this in controlled experiments that keep the graph model and training procedure fixed while replacing the trained auxiliary encoder with an untrained network of the same architecture. Markedly, on QM9 and Alchemy, random SMILES targets reduce mean raw-feature linear-probe error by 5.0% and 3.0% relative to graph-only training, respectively. The effect varies by settings: trained targets matter more for 3D supervision on PCQM4Mv2, for longer training on QMugs, and under GraphMVP fine-tuning. Additional controls show that random graph encoders and fixed random codes can also help, while simple representation expansion does not explain the gains. Together, these results show that cross-modal improvements can arise even when the second encoder has learned no molecular information. A frozen random target should therefore be included as a baseline before attributing such gains to transferred cross-modal content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.