acceptodds
Under review as a conference paper at ICLR 2027

GraphBTE: Graph Byte Triple Encoding

Abstract

Molecular property prediction commonly uses graph-based GNNs or SMILES-based Transformers, where SMILES tokenization captures molecular substructures beyond individual atoms. However, BPE-based SMILES tokenization has two problems regarding invariance under SMILES enumeration: SMILES-enumeration dependence and inconsistent subgraph tokenization. Direct tokenization of molecular graphs addresses these limitations but leads to two additional problems: connectivity uncertainty in pair-based merging and loss of atom-level positions. To address these limitations, we propose Graph Byte Triple Encoding (GraphBTE), which represents each merge rule using two adjacent tokens and the canonical representation of their merged subgraph, while incorporating atom-level positional information through spectral positional encoding (SPE). Across seven MoleculeNet classification datasets, GraphBTE outperforms GraphBPE with a GNN and achieves higher average ROC-AUC than SMILES BPE under the same Transformer architecture. Further analyses show that GraphBTE is invariant to SMILES enumeration, consistently tokenizes shared subgraphs in the analyzed examples, and distinguishes different merged subgraphs formed from the same token pair. SPE further improves prediction performance, and the best-performing vocabulary size is strongly associated with molecular size.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.