Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
Abstract
Deep Research Agents (DRAs) show great potential in assisting humans with complex research tasks. Yet existing DRA evaluations are limited by scalability, as models iterate rapidly, benchmarks require continuous updates, while high-quality references and criteria depend on substantial human effort that cannot scale at low cost. Existing benchmarks have yet to achieve both reference quality and scalability simultaneously. To bridge this gap, we introduce Wiki Live Challenge (WLC), a scalable, live benchmark that leverages the newest Wikipedia Good Articles (GAs) as expert-level references. The Wikipedia community continuously produces expert-reviewed GAs, enabling WLC to acquire high-quality new tasks at minimal cost. Grounded in Wikipedia's GA criteria, we collect 100 recent GAs and propose Wiki Eval, a comprehensive evaluation framework comprising Wiki Writing, a fine-grained writing assessment with 39 criteria, and Wiki Fact, rigorous metrics for factual verifiability. Extensive experiments on various DRA systems demonstrate a significant gap between current DRAs and human expert-level Wikipedia articles.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.