Assess, Answer, Audit: A Point-in-Time Benchmark for Evidence-Grounded Financial Agents From Earnings Calls
Abstract
When assisting financial analysts, agents may engage in multi-turn interactions as analysts clarify and refine their requests. In this scenario, agents are required to interpret evolving requests, make decisions based on available evidence, and evaluate the reliability of answers. Yet existing benchmarks rarely assess these capabilities jointly in real multi-turn financial interactions. Inspired by earningscall dialogue, we introduce AAA-FIN (Assess–Answer–Audit for Finance), a benchmark built on real analyst–management exchanges, preceding Q&A, and the latest pre-call SEC 10-K. Three linked tasks, ASSESS, ANSWER, and AUDIT, evaluate question interpretation and evidence-based response decisions, complete and grounded answer generation, and judgments of historical answers’ evidence support and responsiveness. Together with complementary metrics, these tasks provide a multidimensional profile of financial agents’ strengths and weaknesses. We develop a reusable annotation pipeline combining LLM proposals, critic feedback, and deterministic validation, followed by review and adjudication by three domain experts. Evaluation of five API-accessible and open-weight models reveals uneven strengths across tasks, with no model consistently leading across the evaluated capabilities. Persistent gaps in evidence assessment, answer completeness, and citation coverage identify priorities for improving financial agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.