acceptodds
Under review as a conference paper at ICLR 2027

LifeSearch: Benchmarking Deep Research Agents for Real-World Life and Productivity

Abstract

Deep research (DR) agents address open-ended information needs through multi-step reasoning, long-horizon web search, and evidence synthesis. However, existing benchmarks primarily focus on academic research or synthetic tasks, leaving DR agents’ capabilities in real-world life and professional productivity settings insufficiently evaluated. To bridge this gap, we introduce LifeSearch, a challenging benchmark comprising 203 Chinese and English tasks derived from real-user query seeds across 7 concrete life and productivity domains. Specially, we spend totally more than 4,500 hours of human annotation to construct over 4,000 expert-written, fine-grained rubrics, covering both report-content criteria for coverage and accuracy, as well as concrete criteria specifying values, units, dates, and URLs. We conduct extensive experiments on 12 advanced DR agents under a unified evaluation harness, systematically analyzing their performance patterns and failure modes. We find that even the best-performing frontier agent, GPT-5.6 Sol, achieves a score of only 63.4%, revealing the potential improvement on realistic life and professional DR tasks. We compare candidate LLM judges against human rubric labels to select the grading configuration. A human ranking study shows strong alignment between aggregate benchmark scores and human judgments, with Spearman's and Kendall's . Overall, LifeSearch provides a rigorous evaluation framework and diagnostic foundation for advancing DR agents toward real-world life and productivity applications.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.