MolToken: Scaling a Molecular Foundation Model with Unified Graph Tokenization across Chemical and Biological Space
Abstract
A central goal of molecular foundation modeling is to jointly model chemical and biological space. However, existing molecular foundation models often retain class-specific representations and encoding schemes even in joint modeling, while scaling behavior under a unified chemical representation remains underexplored. We introduce MolToken, a Molecular foundation model with unified graph Tokenization across chemical and biological space. MolToken maps small molecules, peptides, proteins, DNA, and RNA into a unified atom-level graph space. Strictly reversible Frequency Dual Depth First Search (FDDFS) serialization and Byte Pair Encoding (BPE) yield a unified vocabulary for joint pre-training with a single autoregressive Transformer. BPE reduces the total length of FDDFS sequences by a factor of approximately 71.8, making atom-level modeling more tractable. We systematically compare joint and single-domain pre-training across model sizes and data scales. Within the evaluated range, compute-optimal model and data scales follow power laws, while optimal loss decreases approximately linearly with log compute. Out-of-sample experiments support loss prediction at substantially larger compute. Cross-domain comparisons reveal domain-dependent benefits and trade-offs of joint training and suggest that single-domain scaling cannot substitute for cross-domain data coverage. Downstream fine-tuning further demonstrates the adaptability of the unified pretrained backbone to molecular generation and design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.