acceptodds
Under review as a conference paper at ICLR 2027

Is It a Ranking Failure? A Same-Budget Action-Ordering Audit for Verifier-Guided Logical Reasoning

Abstract

When a verifier-guided reasoning system fails within a fixed budget, can changing only the order of its actions certify the correct answer? We introduce a same-budget audit that answers this question while holding the rules, available actions, verifier, and stopping rule fixed. For finite grounded Horn-style forward chaining, shortest certifying distances define the fraction of instances with some successful ordering within budget. Offline search and saturation-cost arguments determine this ceiling, with explicit bounds when resource-capped search leaves cases unresolved. On 21,677 instances from a restricted ProofWriter test subset and generated PrOntoQA partitions, all 4,988 selected-policy failures at an eight-proposal budget require saturation: the stopping rule waits until every derivable fact has been obtained. Their exact costs are 9–19 proposals, so reordering cannot repair any of these failures: the selected policies attain the ceiling. Lower-budget audits also identify repairable failures, including one ProofWriter instance missed by all thirteen tested policies. Broader data scope, changed rule order, and alternative stopping rules reveal the limits of the zero-gap finding. The audit distinguishes failures caused by action ordering from those constrained by the fixed certification protocol, providing a basis for deciding which component to improve.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.