acceptodds
Under review as a conference paper at ICLR 2027

SIGMA: Semantic Identifier Grouping for Molecular Autoregression

Abstract

Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization. Equivalent serializations can induce different internal states along a shared molecular continuation. Randomized strings broaden exposure, but do not explicitly supervise correspondence between those states. We introduce SIGMA, a dense suffix-position objective built from chemically certified same-suffix triplets: two equivalent histories, one non-equivalent history, and a shared suffix. SIGMA aligns corresponding pre-token hidden states along the continuation and enforces a finite similarity margin over the non-equivalent history, leaving the language-model objective, decoder, and inference procedure unchanged. We compare SIGMA with canonical training, randomized-serialization training, and last-token alignment across four datasets under SMILES and SELFIES. SIGMA lowers test-reference Frechet ChemNet Distance against every control in all three training seeds on all four SELFIES datasets and on QM9 and ZINC under SMILES. Position-wise analyses show improved state correspondence, chemical discrimination, and top-1 next-token agreement while preserving between-molecule information. Full-corpus ZINC ablations under both representations identify correct state correspondence, rather than extra computation alone, as an effective ingredient. Beyond generation, a whole-molecule SIGMA extension improves mean predictive performance on all six molecular property benchmarks and reduces serialization sensitivity on every task relative to a matched randomized-serialization control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.