MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents
Abstract
Large language models (LLMs) have shown strong general reasoning capabilities, yet molecular design remains challenging because it requires precise structural reasoning, specialized computational tools, and iterative search over a large chemical space. We present MolLingo, a multi-agent framework that addresses these challenges through chemically informed representations, tool-grounded reasoning, and coordinated optimization. We introduce BRICS-based Fragment Enumeration (BFE), which combines chemically valid BRICS fragmentation with corpus-level fragment frequencies to represent molecules as recurring chemical blocks. Encoding these blocks with both SMILES and common chemical names exposes meaningful structural units while preserving molecular connectivity. Building on this molecular representation, MolLingo integrates a Chemist Agent, a Literature Agent, an Orchestrator, and shared memory, together with tools for ADMET prediction, molecular docking, and binding retrieval. For property optimization, across GPT-5.4, Claude-4.6-Sonnet, and Gemini-3-Pro, block-based input improves the joint objective of property improvement and structural similarity over raw SMILES under matched LLM backbones. For binding optimization, incorporating docking-derived structural context increases mean docking-score improvement from 2.4% with direct prompting to 10.4% with GPT-5.4. Iterative experiments further show that memory improves structural preservation and reduces redundant edits, while adaptive search allocation improves optimization as the candidate budget increases. Together, these results show that molecular representation, structural context, and agentic coordination provide complementary signals for LLM-guided molecular design. Code is available at https://anonymous.4open.science/status/MolLingo-7450.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.