acceptodds
Under review as a conference paper at ICLR 2027

Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?

Abstract

An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record *evidentiary* when a reader who was not there can check it without trusting the writer. Across sixteen deployed frameworks, none writes one in full. Hearsay examines the record, not the task: five harnesses run fourteen tasks, three blinded LLM examiners and a human panel read the records, and every excerpt an examiner quotes is checked mechanically for who wrote it. Two findings follow. First, the record lets a reader name the fault but not prove how the run went. Examiners name the right fault in 74 to 91% of 140 runs by a majority of three model graders, and the fault can be proved, but only from two files the benchmark adds, the failing test's output and the final diff; for what happened in between, fewer than one citation in ten lands on anything the harness did not write, and when we delete, rewrite or fabricate entries in copies of the records, the examiner with the fewest false alarms catches half of them. Second, the remedy is a second author, not a stronger seal on the first. An append-only log of what passes between harness and model, kept outside the harness, is read against the record in both directions: a query finds what the log saw and the record lacks, a reverse check what the record holds and the log never saw. Together they report all 28 omissions and fabrications we made a harness commit as it ran, where a hash chain over the harness's own record passes all 28, since it hashes whatever the harness wrote. On 115 records we falsified by editing copies they report 98, each scored against a clean copy of its run with rules written on those pairs; held out on another vendor's records, the reverse check catches 29 of 34 fabrications. Handed the log, examiners keep their fault verdicts but rest more of their citations on what the harness did not write. What makes a record evidence is who writes it, not what is captured.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.