acceptodds
Under review as a conference paper at ICLR 2027

Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check

Abstract

When a multi-agent pipeline fails, an LLM auditor is asked which step let the fault in. It reads the agents' reports, whose suspected-origin field can state a conclusion ("suspected origin: step 3") beside their observations. Holding the case evidence fixed, we intervene on that conclusion alone (chain agents gpt-4.1-mini and claude-haiku-4-5, auditor qwen3-30b-a3b-instruct). The auditor relays the conclusion rather than checking it. In a registered single-run re-measurement on a pinned provider, with the reports, the conclusion and the documentation each step received all in its prompt, it finds the true origin in only 4.0% (GPT) and 2.9% (Claude) of the episodes where the agents' proposal is wrong, against 59.7% and 67.5% from the same documentation alone: the evidence goes unused on GPT chains, largely unused on Claude chains. From the reports alone, deleting only the conclusion raises accuracy where the proposal is wrong by 38.3 pp (GPT) and 24.0 pp (Claude), a gain that also appears with gpt-5, claude-sonnet-5 and gemini-2.5-flash-lite as auditors and in a second domain. Removal has a price: where the conclusion is right it costs 18.3 and 10.5 pp, less consistently across auditors, both directions holding after Holm correction whether or not naming the step just before the origin is scored correct. Without the conclusion the auditor still trails the documentation alone (post hoc). Whoever writes the field steers the audit (pre-specified). A random wrong step planted in it is adopted in 81.9% (GPT) and 64.1% (Claude) of the episodes tested, nearly as often as the agents' own proposals (0.83 of their gain over the no-field rate), and causally impossible plants still in 19.4% and 29.2% (descriptive). An instruction not to use the field removes only 24% (GPT) and 37% (Claude) of the auditor's dependence on it and costs 7 to 9 pp where the proposal is right. The mechanism reading, by a rule frozen before collection, is mixed, with a measurable numeric-anchoring component. Whether deletion pays depends on how the origin is scored, and our pre-registered responsibility-framing hypothesis is not supported.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.