acceptodds
Under review as a conference paper at ICLR 2027

MolToken: Scaling a Molecular Foundation Model with Unified Graph Tokenization across Chemical and Biological Space

Abstract

A central goal of molecular foundation modeling is to jointly model chemical and biological space. However, existing molecular foundation models often retain class-specific representations and encoding schemes even in joint modeling, while scaling behavior under a unified chemical representation remains underexplored. We introduce MolToken, a Molecular foundation model with unified graph Tokenization across chemical and biological space. MolToken maps small molecules, peptides, proteins, DNA, and RNA into a unified atom-level graph space. Strictly reversible Frequency Dual Depth First Search (FDDFS) serialization and Byte Pair Encoding (BPE) yield a unified vocabulary for joint pre-training with a single autoregressive Transformer. BPE reduces the total length of FDDFS sequences by a factor of approximately 71.8, making atom-level modeling more tractable. We systematically compare joint and single-domain pre-training across model sizes and data scales. Within the evaluated range, compute-optimal model and data scales follow power laws, while optimal loss decreases approximately linearly with log compute. Out-of-sample experiments support loss prediction at substantially larger compute. Cross-domain comparisons reveal domain-dependent benefits and trade-offs of joint training and suggest that single-domain scaling cannot substitute for cross-domain data coverage. Downstream fine-tuning further demonstrates the adaptability of the unified pretrained backbone to molecular generation and design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.