acceptodds
Under review as a conference paper at ICLR 2027

Faultless Misreading: Diagnosing Professional Agents Under Incomplete Spreadsheet Requests

Abstract

Large Language Models (LLMs) are increasingly deployed to automate professional work, yet real-world user requests in these domains are inherently underspecified. Evaluating an agent's ability to clarify ambiguities and align with human intent is therefore critical. We study this challenge in spreadsheet manipulation, the everyday task of editing workbooks from written instructions. We introduce PIVOSM (Paired Instructions with Verifiable Open-decision Spreadsheet Manipulation), a benchmark of 434 matched instruction pairs over spreadsheet workbooks. It isolates decision-critical degrees of freedom across three axes of underspecification (Missing Key Information, Multiple Reference and Erroneous Information). Powered by double-layer programmatic rubrics, it provides fully automated scoring that disentangles code execution from human intent alignment. Our evaluation of nine frontier models across explicit, implicit, and interactive settings shows that underspecifying one degree of freedom reduces intent accuracy by , while execution accuracy moves by only , and that letting agents ask closes only half of the resulting intent gap on average. This suggests that while execution capabilities remain stable, recovering unstated human intent via proactive clarification represents the true bottleneck for professional agents operating under realistic underspecification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.