TaxLab: An AI Tax-Filing Benchmark Built on Field-Level Operationalization of Tax Law
Abstract
AI agents are increasingly capable of completing complex, multi-step tasks, making high-stakes applications, such as automated tax preparation, a promising direction. Evaluating tax-filing agents, however, remains difficult: it is hard to determine what range of tax situations a benchmark actually covers, how diverse those situations are, and where an agent's mistakes originate when a completed return is incorrect. We introduce TaxLab, a benchmark for interactive U.S. federal individual tax filing built on field-level operationalization of tax law. We construct TaxLab-Rule, an executable, source-grounded rule system comprising 3,142 rules, each linking a specific reporting field to the official sources that determine its treatment. This representation provides a common foundation for constructing gold returns, measuring benchmark coverage and diversity, and tracing filing errors back to the rules and sources involved. Using this system, we construct 55 CPA-reviewed tax-year-2025 scenarios and an interactive environment in which agents must identify missing taxpayer information, query a simulated taxpayer, consult offline IRS guidance, and complete the required tax forms. We evaluate nine models and find that tax preparation remains challenging: the strongest fully evaluated model, Claude Opus 5.5, achieves only 64.0% perfect filing accuracy despite substantially higher field-level accuracy. Our error diagnosis further reveals distinct issue-level failure patterns across models, showing that even strong models remain disproportionately error-prone on specific tax treatments. Together, these results show that substantial progress is still needed before AI agents can reliably prepare complete tax returns, and that improving aggregate accuracy alone is insufficient without addressing the specific tax treatments that drive filing failures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.