ANOTHER SAMPLE OR A BETTER VERIFIER? WHAT INFERENCE BUDGET BUYS FOR CONSISTENT AGENTS
Abstract
An agent that handles refunds is useful only if it gets them right every time, not most of the time. Benchmarks call this consistency and measure it as passk, the chance that all k attempts at a task succeed. We ask what a fixed byte budget per conversation should buy to raise it, a sample, a verifier call, a vote, or a retry, and we find a rule. At the budgets operators run, how cheaply a candidate is selected matters ten times more than how accurately; we call this the cost-over-accuracy rule. The evidence comes from replaying 6,428 released τ-bench and τ 2-bench trajectories through an evaluator that scores any allocation policy exactly, running no model. Across four frontier models and three domains, a single attempt succeeds 56.3% of the time and all four attempts 41.4%. Buying operations rather than samples raises pass4 by +1.76 points at equal cost. A perfect verifier that must read every candidate reaches 43.6%, no better than a vote, because reading costs as much as sampling and cuts the draws in half; the same verifier made free reaches 53.2%, and a proposition predicts where the two cross. A verifier that reads only a candidate’s write actions, 2% of the bytes, follows the rule and reaches 47.7%. Two expectations written down before the study were refuted. Optimizing for consistency changed no allocation, and a learned controller matched but did not beat a tuned rule.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.