acceptodds
Under review as a conference paper at ICLR 2027

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning

Abstract

Spreadsheet systems such as Microsoft Excel and Google Sheets are central to modern data workflows, and AI agents that operate them are a promising research direction. Existing spreadsheet agents mostly prompt general-purpose LLMs, which works for simple operations but struggles with the multi-step workflows that dominate real-world use. We introduce Spreadsheet-RL, a reinforcement learning (RL) framework for training specialized spreadsheet agents in a realistic Microsoft Excel environment. It comprises an automated data agent that collects paired initial–final spreadsheets from online forums; Spreadsheet Gym, a multi-turn Excel environment with a spreadsheet-native tool harness and an asynchronous Excel-based verifier; and Domain-Spreadsheet, a new benchmark of 1,660 professional tasks covering finance, supply chain, human resources, sales, and real estate. On SpreadsheetBench, Spreadsheet-RL raises Qwen3-4B-Thinking-2507 from 12.0% to 23.4% Pass@1, with gains from both the spreadsheet-native harness and RL post-training; the gains hold at a larger model size on Qwen3-8B (16.8% to 22.3%) and generalize to Domain-Spreadsheet (8.4% to 17.2%). A human audit of 100 training and 100 Domain-Spreadsheet tasks finds majority-valid oracles for 84% and 98% of tasks, and trajectory analysis shows that RL mainly improves interaction reliability, cutting tool-execution errors from 48.9% to 7.4% of calls. These results highlight Spreadsheet-RL's strong potential for generalization and real-world adoption in spreadsheet automation, and broadly, its promise for advancing LLM-based interactions with data interfaces in everyday work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.