Assurance, Not Just Accuracy: Auditing LLM Reasoning with Audit-Risk Control
Abstract
Benchmarks report how accurate a large language model (LLM) is on average; users of a single answer need to know how much they can trust this output. We transplant the institutional design of financial auditing into LLM quality control and study the resulting system, Ara (Audit-Risk control for Assurance). Ara converts a target misstatement risk into per-claim detection-risk budgets via the audit-risk model (inherent, control, and detection risk), purchases evidence under a token budget by risk-weighted knapsack, ranks evidence by independence (CPU recomputation over isolated LLM procedures over non-isolated ones), and issues per-output audit opinions. The deliverable is an empirical dual-caliber certificate, conditional on the realized certified subset of the measured distribution: a one-sided Clopper–Pearson 95% upper bound on residual error over the extracted claim universe, at a stated coverage; it is not a per-answer or future-sample guarantee. is a target, not a promise, and we falsify it openly: across twelve trackdomain settings is falsified in three, while 0.05 and 0.10 hold wherever the pass is not vacuous. On GSM8K with Qwen2.5-7B, Ara certifies 89.0% of answers at 3.8% residual error (raw: 10.9%). Auditing the frontier API model DeepSeek V4 Pro certifies 93.4% of GSM8K-slice answers at 2.1% risk (CP95 3.6%) and 57.7% of stress answers at zero error as-run; a deterministic rejudging recovers coverage to 96.2%/78.2%. Root-causing attributes 98.7% of adverse evidence to our own extraction layer, dominated by a single Unicode minus-sign artifact, not arithmetic inability. A sealed reproduction on a second, weaker deployment (Tencent Hunyuan Hy3, reasoning disabled, ¥15.43) re-confirms the certificate on GSM8K, does not confirm it on stress, and finds evidence-channel value generator-dependent, not universal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.