TadA-Bio: A Replayable Biomedical AI-Agent Benchmark for Budgeted Directed Evolution
Abstract
Scientific agents in biological discovery campaigns must coordinate sequential operations: choosing which branch to open, how much evidence to collect, which candidates to validate, and when to stop under finite time and budget. Existing scientific evaluations in protein engineering largely emphasize sequence prediction or candidate ranking, leaving the role of an agent as the operator of a budgeted, long-horizon campaign underexplored. We introduce TadA-Bio, a replayable benchmark for evaluating such decision policies. It is built from measured sequences, activities, and provenance from a historical 31-round TadA directed-evolution campaign, but adds a controlled replay layer: each episode exposes a partial public state, hides future candidates and activities, and uses fixed benchmark costs, times, screening noise, and reward rules. An agent can prepare seeds, open libraries, screen, confirm hits, acquire BioSkills, or stop; terminal cash is the primary outcome. Across heuristic policies, three LLM-agent harnesses (Codex, OpenCode, and Claude Code), budget sweeps, and BioSkill settings, the benchmark separates productive operation from cash preservation, wasteful exploration, runtime failures, and unproductive skill purchase. TadA-Bio is a playable, discriminative, and diagnostic offline benchmark, providing a grounded example and potential foundation for future automated biological laboratories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.