XLGym: API-Grounded Task Synthesis for Training and Evaluating Spreadsheet Agents
Abstract
Spreadsheet agents act on tables, charts, PivotTables and conditional formats, yet existing benchmarks are dominated by formula, computation and domain-oriented queries, which frontier models are getting increasingly good at. We introduce XLGym, a pipeline that turns Excel API documentation and real workbooks into hard tasks with reference solutions. Each task is grounded in a specific API member, verified by execution in Excel, and hardened until frontier models no longer solve it reliably. Because tasks are seeded per API member and per artifact, coverage and difficulty are set by construction. With XLGym we build XLBench, a 190-task benchmark that covers all eight major artifact types and gives every frontier model its lowest score among the benchmarks we evaluate (Opus 5: 52.8%, GPT-5.6: 35.4%), and XLTrain, a 20.6k-task corpus spanning 1,287 API members. Fine-tuning Qwen3.6-27B on XLTrain yields XLSage-27B, which raises the pass rate from 1.8% to 31.9% on XLBench, from 3.1% to 31.6% on WTM-Bench and from 2.5% to 22.2% on SpreadsheetBench. Because XLGym can regenerate tasks against newer models, XLBench can continually evolve.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.