StateEdit-Bench: From Local Edits to Stateful Presentation Workflows
Abstract
Presentation editing requires interpreting later requests in light of previously established objects, decisions, constraints, and edit boundaries. Yet existing evaluations primarily measure fulfillment of the current request or properties of the final output, leaving the preservation and selective revision of accepted state weakly observed. We introduce STATEEDIT-BENCH, a 126-unit benchmark suite with 218 goal instances on 58 distinct input decks: 70 single-turn local edits, 10 continuation or template extensions, and 46 three-goal stateful sessions (138 sequential goals). Each stateful goal advances the committed artifact and is evaluated with an edit contract covering current postconditions, retained constraints, artifact validity, and authorization-aware preservation. We also design SE-PPT-EDIT, a presentation-editing skill that combines explicit state tracking, scope-aware planning, and artifact verification. In paired no-skill/full-skill runs across three models, the equal-model descriptive macro changes from 69.8% to 75.8% (+6.0 pp), with the largest gain on continuation and template filling, followed by local editing. On stateful multi-turn sessions, however, the macro changes by only +2.2 pp, and the model-level changes are not statistically significant. These findings position StateEdit-Bench as a diagnostic benchmark: it exposes contract-state failures that current-request evaluation can miss, while showing that the evaluated procedural deployments do not yet reliably improve multi-turn state maintenance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.