SIGMA: Semantic Identifier Grouping for Molecular Autoregression
Abstract
Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization. Equivalent serializations can induce different internal states along a shared molecular continuation. Randomized strings broaden exposure, but do not explicitly supervise correspondence between those states. We introduce SIGMA, a dense suffix-position objective built from chemically certified same-suffix triplets: two equivalent histories, one non-equivalent history, and a shared suffix. SIGMA aligns corresponding pre-token hidden states along the continuation and enforces a finite similarity margin over the non-equivalent history, leaving the language-model objective, decoder, and inference procedure unchanged. We compare SIGMA with canonical training, randomized-serialization training, and last-token alignment across four datasets under SMILES and SELFIES. SIGMA lowers test-reference Frechet ChemNet Distance against every control in all three training seeds on all four SELFIES datasets and on QM9 and ZINC under SMILES. Position-wise analyses show improved state correspondence, chemical discrimination, and top-1 next-token agreement while preserving between-molecule information. Full-corpus ZINC ablations under both representations identify correct state correspondence, rather than extra computation alone, as an effective ingredient. Beyond generation, a whole-molecule SIGMA extension improves mean predictive performance on all six molecular property benchmarks and reduces serialization sensitivity on every task relative to a matched randomized-serialization control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.