acceptodds
Under review as a conference paper at ICLR 2027

Learning what to preserve: Grammar constrained search across scientific domains

Abstract

Scientific strings encode structure directly in their notation. A chemical formula such as specifies elemental identities and stoichiometric quantities, while a protein mutation such as M105K specifies residues and a positional coordinate. Yet pretrained models typically process such inputs using tokenizers learned from corpus statistics rather than from the structure of the notation itself. Existing solutions often introduce domain-specific parsers, vocabularies, or tokenizers, requiring a new representation strategy for each scientific domain. We ask whether a shared input representation rule can instead be discovered from downstream tasks across domains with different structural requirements. We introduce GCMS, a grammar-constrained multi-domain search framework that uses a policy gradient search over interpretable input rewrites. Joint search on crystal-system classification and protein mutation fitness discovers a common representation without an element inventory, residue classifier, or domain-specific parser. The learned rule transfers across unseen tasks and notation systems, including held-out materials properties, additional protein assays, modified peptides, and genomic variants. GCMS reduces band-gap MAE from to relative to WordPiece and reaches Spearman on the held-out ProteinGym BLAT assay. These results show that transferable structure in scientific notation can be discovered from task supervision rather than engineered separately for each domain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.