Learning Reliable Skills from Sparse Online Feedback for Cold-Start LLM-Based Molecular Optimization
Abstract
Molecular optimization seeks drug-like candidates with strong target-specific properties under small budgets of costly evaluations. LLM-based optimizers are attractive because they propose molecules from natural-language objectives and adapt in context without training a task-specific generator. Yet methods that endow an optimizer with reusable skills typically require large offline trajectories, either to train a policy or to build a retrieved skill library — resources a cold-start target does not have. While simpler cold-start methods such as OPRO present candidates as a score-ranked list of independent molecules, preserving scores but discarding the skills latent in the online history. We propose OS-DEV (Online Skill Distillation with Evidence-Aware Verification), a training-free feedback layer that turns each evaluation into a best-referenced trace and periodically distills improving and degrading traces into target-specific proposed skills. Two independent analysts extract complementary hypotheses about what to pursue and avoid, and a role-independent verifier re-examines each proposed skill against the supporting and contradicting observations accumulated in the current run, attaching an interpretable high/medium/low confidence judgment before the resulting verified skill state guides subsequent generation. On all 58 DOCKSTRING targets with 200 oracle evaluations per target and a shared Qwen3.5-4B generator, OS-DEV reaches an overall score of 0.8487, improving over ACE (0.8254) by 0.0233 and over an agentic reference combining Claude Code with DeepSeek-V4-Pro (0.8270) by 0.0217, and leads on 30 of 58 targets. Ablations isolate the contributions of the trace display, the dual analysts, and the evidence-aware verifier, and confirm that organizing sparse feedback into reusable, evidence-linked skills improves LLM-based molecular optimization without retraining the proposal model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.