acceptodds
Under review as a conference paper at ICLR 2027

AUTOMATICALLY BUILDING MOLECULAR PROPERTY PREDICTION MODELS WITH AN EXPERT-GUIDED AI AGENT

Abstract

Molecular property prediction guides compound prioritisation in drug discovery, but turning domain expertise into an effective model takes machine-learning engineering and repeated cycles of design, implementation, training and evaluation. We present a hierarchical agent pipeline that automates this workflow. Given a task description, the data and optional written expert guidance, it runs an agent loop that designs a model, writes and debugs the code, trains and tunes it, evaluates the result and, when the results call for it, revises the design, then selects and delivers a final model, so a domain expert steers development in prose rather than code. A controlling language-model agent assigns these steps to specialised sub-agents and decides each next step from their reports and training results; the pipeline code is the same for every task. We evaluate it retrospectively on the OpenADMET pregnane X receptor (PXR) induction challenge, a recent blind competition, using guidance we wrote from published challenge reports and earlier runs of this pipeline after the competition ended. Across three runs, the guided pipeline achieves a mean relative absolute error of 0.561, compared with 0.627 without guidance; the published winning score of 0.563, obtained using proprietary data, provides an external reference. On two further benchmarks, hepatocyte clearance and plasma protein binding, the same pipeline with task-specific guidance scores better on average than the published leaderboard bests. Given guidance that names no model, the pipeline's run fell within the range of three AIDE runs and scored better than three MLEvolve runs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.