Investment Research Agent Benchmark: Grounding Agent Evaluation in Real-World Tasks and Live Information
Abstract
As AI agents enter investment research workflows, their evaluation must reflect diverse professional objectives and changing evidence. This poses three challenges: task grounding, scoring alignment, and reference maintenance. To address these challenges, we introduce the Investment Research Agent Benchmark (IRAB), comprising 101 tasks adapted from a corpus of 252,015 practitioner requests. Focused primarily on Chinese markets, these tasks span 12 research intents, ranging from industry-chain transmission analysis to stock and fund screening. For each task, IRAB pairs a task-specific rubric with dated reference evidence maintained under explicit refresh policies. Each rubric selects and weights applicable criteria from an inventory of 27 reusable, expert-derived criteria. A multiplicative scoring rule then combines data capability, answer quality, and reliability, so that detected factual failures reduce the overall score. Using IRAB, we evaluate 58 hosted configurations of 42 base models. Descriptive comparisons show higher data timeliness and source authority in newer release cohorts, while case analyses reveal persistent factual and verification failures. Repeated-run analysis further identifies variation in cited sources, underscoring the need for evidence preservation. This work thus contributes a practitioner-grounded benchmark, a reliability-aware evaluation framework, and an empirical characterization of the capabilities and limitations of investment research agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.