PPTXEdit-Bench: A Diagnostic Benchmark for Language-Guided PowerPoint Editing
Abstract
Language-driven presentation editing requires a system to map user-facing expressions such as “the card on the left” to native PPTX objects, apply the requested change, and preserve unrelated content. Existing evaluations do not jointly represent editing intent, object roles, and outcome states, so an end-to-end failure cannot reveal whether a system misunderstood the request, selected the wrong objects, or executed the edit incorrectly. We introduce PPTXEdit-Bench, containing 22,100 verifiable editing instances over 4,420 source slides. Each instance pairs a natural-language instruction with a structured edit specification, role-aware object bindings, and state checks for target satisfaction and non-target preservation, covering referents from individual elements to composite entities and predicate-defined sets. We further develop SlideDiag, an inspectable intent–ground–execute pipeline. Gold interventions on intent and grounding measure the performance recovered by correcting upstream predictions and isolate execution errors that remain under fully specified edit plans. We evaluate general-purpose coding agents, one-shot structured editors, and SlideDiag under a unified protocol. PPTXEdit-Bench therefore supports both end-to-end comparison and diagnosis of where precise presentation editing fails.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.