When Is a Cheap LLM Audit Decisive? Prospective Forecasts and Measured Costs
Abstract
The value of efficient evaluation lies in the decisions its evidence supports. We introduce a prospective audit protocol that separates false-assertion control, the probability of a correct definitive decision, and the cost of acquiring model outputs. Before observing new target outcomes, a fixed planner chooses among whole-question sampling, a paid target pilot, partial trajectories, and complete execution. On two new question pools shared by Qwen3 and Ministral3, the selected CommonsenseQA audits detect excessive risk with probabilities 97.51% and 98.80% at 47.5% and 40.1% of full-evaluation cost when source logs already exist. OpenBookQA audits cost 51.5% and 58.1% but confirm the acceptable pools with probabilities 63.97% and 26.83%. Constructing the source logs used by the planner raises its four evaluation-cost ratios to 1.25–1.76. All four prospective uniform-audit forecasts underestimate decision probability. Exact sample-size curves, same-observation inference controls, and an algorithmic CollabEval comparison distinguish the roles of sampling, inference, and information cost. These measurements connect evaluation efficiency to decision readiness under fixed finite-pool designs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.