PM Bench: Verifiable Clarification Benchmark for Coding Agents
Abstract
As coding agents take on increasingly autonomous software engineering, a critical capability emerges: reliably and consistently resolving incomplete and contradictory specifications. We argue that this is a fundamental capability that requires objective ground truth that is not a matter of judgment. We introduce PM Bench, a benchmark that measures two complementary capabilities: detection, a metric for asking the correct clarification that resolves the issues injected into a complete specification, and restraint, a metric for not manufacturing clarification questions about a complete specification. Our central contribution is a fully automated, formalism-first construction in which a computer science formalism, rather than natural language, is the source of truth across both task generation and verification. We begin from a certified, well-specified statechart, plant controlled defects (e.g., an unreachable state, conflicting transitions, a counter overflow), and render the result into a Product Requirements Document that reads as complete prose. Because the answer key is the planted defects, scoring is verifiable: a hidden judge checks each clarifying question against the formal oracle facts. Evaluating eleven models, we find that detection is near the ceiling for frontier models while restraint is not: models are separated almost entirely by their false-alarm rate on complete specifications. We further find that the choice of scoring metric reorders the middle of the roster. A paired pilot shows that re-rendering the same certified chart as human-style prose leaves detection unchanged but erodes restraint in frontier models. The method is demonstrated on controller statecharts but generalizes to any domain expressible as a formal specification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.