Faultless Misreading: Diagnosing Professional Agents Under Incomplete Spreadsheet Requests
Abstract
Large Language Models (LLMs) are increasingly deployed to automate professional work, yet real-world user requests in these domains are inherently underspecified. Evaluating an agent's ability to clarify ambiguities and align with human intent is therefore critical. We study this challenge in spreadsheet manipulation, the everyday task of editing workbooks from written instructions. We introduce PIVOSM (Paired Instructions with Verifiable Open-decision Spreadsheet Manipulation), a benchmark of 434 matched instruction pairs over spreadsheet workbooks. It isolates decision-critical degrees of freedom across three axes of underspecification (Missing Key Information, Multiple Reference and Erroneous Information). Powered by double-layer programmatic rubrics, it provides fully automated scoring that disentangles code execution from human intent alignment. Our evaluation of nine frontier models across explicit, implicit, and interactive settings shows that underspecifying one degree of freedom reduces intent accuracy by , while execution accuracy moves by only , and that letting agents ask closes only half of the resulting intent gap on average. This suggests that while execution capabilities remain stable, recovering unstated human intent via proactive clarification represents the true bottleneck for professional agents operating under realistic underspecification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.