ChemoGrad: Improving Chemical Reasoning through Test-Time Latent Optimization
Abstract
The latent space of chemical LLMs encodes the underlying characteristics and principles of chemical structures, making it possible to generate chemical compounds by reasoning directly within this space. Yet existing latent-reasoning methods for chemical generation lack a clear direction: the latent state is updated by re-forwarding it through the model, without a reward signal or verifier to steer the search. We propose ChemoGrad, which optimizes molecules by a reward-guided search over a frozen model's latent states: it scores each candidate with a property evaluator, nudges the hidden states to perturb the emitted prefix, then re-decodes and keeps a candidate only when the evaluator confirms an improvement. A likelihood-sharpening step proposes candidates around the model's own trace, a verifiable reward selects which to pursue, and a task-oriented initialization points the search in the right chemical direction — shifting task adaptation from training-time parameter learning to inference-time search. On ChemCoTBench, ChemoGrad consistently outperforms the state-of-the-art baseline across all six molecule-optimization tasks, gaining +51% in mean property improvement and +175% on JNK3 activity, with no additional training; it also transfers to editing, reaction prediction, and captioning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.