acceptodds
Under review as a conference paper at ICLR 2027

Where the Thread Breaks: Evaluating LLM Auditors in Long-Horizon Scientific Workflows

Abstract

Advances in large language models (LLMs) have enabled autonomous agents to perform complex, long-horizon tasks. These tasks require reliable output verification, as errors can propagate across steps and compromise the final results. LLM-based auditing provides a practical approach to verifying the output of each step. However, its performance in long-horizon workflows, where later steps depend on earlier outcomes, remains underexplored. We introduce AuditSpan, a benchmark for evaluating LLM-based auditing in long-horizon scientific workflows. Specifically, we (1) examine whether and how the benefits and limitations of auditing individual outputs carry over to long-horizon workflows and (2) investigate challenges arising from interactions among workflow steps, including difficulties in assessing overall scientific validity and tracking error propagation and correction. Our results show that, when auditing extends from individual outputs to long-horizon workflows, appropriate abstention decreases and overconfidence increases. Beyond these amplified limitations, auditors may approve scientifically flawed workflows when most individual steps are correct and may treat errors as resolved even when their effects persist in later outputs. These findings can guide the development of more reliable LLM-based auditors and support progress toward fully automated long-horizon workflows.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.