acceptodds
Under review as a conference paper at ICLR 2027

What, Where, and How to Edit: Benchmarking Instruction Following for Localized Molecular Editing.

Abstract

Large language models are increasingly used to generate and modify drug-like molecules from natural-language instructions, yet existing benchmarks rarely test whether a model follows compositional editing requests precisely. We introduce IF-DrugBench, a benchmark of 2,400 localized small-molecule editing tasks organized along three complementary axes: what to achieve, where to act, and how to use reference information. Each task is released in aligned one-dimensional SMILES and two-dimensional molecular-image forms, grounded in an experimentally reported molecular transformation, and paired one-to-one with machine-readable criteria and automated evaluation tools. The benchmark evaluates molecular validity, localization of substituent or linker edits, chemical-operation fidelity, explicit reference-motif transfer when requested, and property directions relative to the starting molecule. On a balanced 480-task evaluation subset, the evaluated models achieve at most 48.1% task pass rate (TPR) with text input and 51.5% TPR with image-conditioned input. These results show that current models still have substantial room for improvement in reliably following compositional molecular editing instructions and satisfying multiple constraints simultaneously. The benchmark and evaluation tools can be checked at https://anonymous.4open.science/r/ICLR-DrugBench-6D27.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.