ReproBAIT: Can coding agents build their own biological AI tools?
Abstract
Large language models are increasingly competent at leveraging their scientific knowledge and existing computational tools to execute complex workflows. However, little information exists on their ability to build tools which are not readily available. We introduce an evaluation titled ReproBAIT (Reproducing Biological AI Tools) which focuses on measuring the ability of coding agents to design, implement and train biological AI tools from published research, across a range of information access conditions. The coding agent is securely sandboxed without internet access, then receives training data, a fixed time and compute budget, and limited information about the tool to replicate. We score the produced biological AI tool on tool-specific metrics over held-out data. We evaluate across a set of biological AI tools of varying complexities: ProteinMPNN, an inverse protein folding tool, RFdiffusion3, an all-atom generative diffusion tool, and a closed-source molecular design tool. We institute a methodology to check for memorisation, as the open source tools are likely to be included in the pretraining corpus. Across a range of leading closed- and open-source coding agents, we find promising reproduction capability for our selected tools. We find moderate evidence consistent with memorisation, varying across coding agents. These results suggest that tool developers may need to be increasingly vigilant about the information that they release publicly. Public information could undermine a tool’s commercial value or invalidate dual-use managed access strategies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.