acceptodds
Under review as a conference paper at ICLR 2027

ANSER-Bench: A Comprehensive Benchmark for Analytical Search

Abstract

Search and retrieval systems are increasingly used to support complex professional analysis and decision-making, where users seek not merely isolated facts but evidence-grounded conclusions to analytical questions. Meeting these needs requires retrieving heterogeneous evidence, performing multi-step aggregation and reasoning, and producing conclusions that remain traceable to their supporting evidence. Yet existing benchmarks often evaluate retrieval, question answering, and reasoning in isolation, leaving end-to-end analytical search insufficiently tested. We introduce ANSER-Bench,Benchmark is available at https://anonymous.4open.science/r/ANSER-Bench-8947. a bilingual benchmark of 2,235 tasks across law, finance, and scientific research, spanning descriptive, predictive, and prescriptive analytical needs. Each task pairs a query with a domain-specific corpus, a reference conclusion, and critical-evidence annotations, enabling separate evaluation of answer quality, evidence quality, and efficiency. Our evaluation reveals substantial heterogeneity across analytical needs: across 14 task settings, no system leads in more than five, and methods that perform strongly in one setting often fail to transfer to others. This variability highlights task-dependent challenges in evidence acquisition and downstream analytical reasoning, positioning ANSER-Bench as a testbed for diagnosing such failures and advancing reliable, evidence-grounded analytical search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.