acceptodds
Under review as a conference paper at ICLR 2027

XLGym: API-Grounded Task Synthesis for Training and Evaluating Spreadsheet Agents

Abstract

Spreadsheet agents act on tables, charts, PivotTables and conditional formats, yet existing benchmarks are dominated by formula, computation and domain-oriented queries, which frontier models are getting increasingly good at. We introduce XLGym, a pipeline that turns Excel API documentation and real workbooks into hard tasks with reference solutions. Each task is grounded in a specific API member, verified by execution in Excel, and hardened until frontier models no longer solve it reliably. Because tasks are seeded per API member and per artifact, coverage and difficulty are set by construction. With XLGym we build XLBench, a 190-task benchmark that covers all eight major artifact types and gives every frontier model its lowest score among the benchmarks we evaluate (Opus 5: 52.8%, GPT-5.6: 35.4%), and XLTrain, a 20.6k-task corpus spanning 1,287 API members. Fine-tuning Qwen3.6-27B on XLTrain yields XLSage-27B, which raises the pass rate from 1.8% to 31.9% on XLBench, from 3.1% to 31.6% on WTM-Bench and from 2.5% to 22.2% on SpreadsheetBench. Because XLGym can regenerate tasks against newer models, XLBench can continually evolve.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.