Can agents outperform human expert crowdsourcing? Insights from DREAM
Abstract
LLM-powered coding agents are increasingly claimed to be able to perform the work of a human scientist on open-ended, real-world tasks in computational biology. Existing benchmarks with pre-specified answers can only establish whether an agent is correct, but not whether it is competitive with a domain expert who attacked the same open problem. Here, we present a living, extensible benchmark of commercial and open-source AI agents on DREAM (Dialogue on Reverse Engineering Assessment and Methods) Challenges, which are crowdsourced competitions with hidden test data, expert-crafted scoring, and human team leaderboards. Our tasks evaluate end-to-end computational biology workflows against expert teams from crowdsourced challenges. Under guided prompt settings, agents' mean score never surpasses the human best, though individual runs did on 2 tasks. Surprisingly, in some cases where agent received less context, agents outperformed the best human solutions and we did not find statistically significant gains with increased context. Worryingly, in some cases when agents outperform the top historical human solution, careful examination of the traces can reveal nuanced ways of exploiting leaked data attributes, which we exclude from our comparisons. Additionally, we compare three test-time optimization methods for improving agent performance on one of our tasks, juxtaposing approaches that optimize the prompt, the code, and agent swarms in an exploratory case study on one task. Agent swarms show the largest increase in performance, consistent with other test-time scaling findings, though careful guards should be implemented to avoid gaming leaderboards. Finally, we introduce a multi-layered evaluation pipeline combining ground-truth metrics, LLM-as-judge scoring, and human calibration to characterize agent's domain know-how. Our open-source framework and leaderboard aim to support the responsible adoption of AI agents in the life sciences, and we invite the community to submit both agents and challenges.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.