Where the Thread Breaks: Evaluating LLM Auditors in Long-Horizon Scientific Workflows
Abstract
Advances in large language models (LLMs) have enabled autonomous agents to perform complex, long-horizon tasks. These tasks require reliable output verification, as errors can propagate across steps and compromise the final results. LLM-based auditing provides a practical approach to verifying the output of each step. However, its performance in long-horizon workflows, where later steps depend on earlier outcomes, remains underexplored. We introduce AuditSpan, a benchmark for evaluating LLM-based auditing in long-horizon scientific workflows. Specifically, we (1) examine whether and how the benefits and limitations of auditing individual outputs carry over to long-horizon workflows and (2) investigate challenges arising from interactions among workflow steps, including difficulties in assessing overall scientific validity and tracking error propagation and correction. Our results show that, when auditing extends from individual outputs to long-horizon workflows, appropriate abstention decreases and overconfidence increases. Beyond these amplified limitations, auditors may approve scientifically flawed workflows when most individual steps are correct and may treat errors as resolved even when their effects persist in later outputs. These findings can guide the development of more reliable LLM-based auditors and support progress toward fully automated long-horizon workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.