acceptodds
Under review as a conference paper at ICLR 2027

Beyond Observability: Using Execution History to Help Agents Self-Diagnose

Abstract

Coding agents increasingly execute sequences of dependent jobs, where an upstream error can propagate through intermediate artefacts and become visible only much later. Diagnosing such failures requires retrospective fault credit assignment: mapping an incorrect downstream result back to the earlier executed function that introduced it. We investigate whether captured execution evidence - executed functions, intermediate runtime states, and exact artefact hand-offs - helps LLM reviewers localise subtle, data-dependent faults, such as threshold-boundary errors, and snapshot-selection errors, whose effects depend on the values encountered during execution. In a controlled four-job study, reviewers received complete source code and top-level inputs and outputs for the full workflow, with conditions additionally exposing different forms of execution evidence (from source code, inputs/outputs and logs through to execution graph and intermediate artefacts and data states). For the easier fault mechanisms, additional execution evidence produced no comparable improvement in localisation. For the harder faults, reviewers were allowed to inspect additional runtime evidence in stages: first selecting a relevant captured execution group, then opening a specific intermediate artefact. On the hard instance, this process surfaced the correct group and row-level artefact in all six sessions and led to the correct fault localisation in five, whereas extra reasoning calls without any new runtime evidence finished at 0/3. These results suggest a broader use for execution graphs beyond observability and governance. An agent could use the execution graph to trace a downstream failure back through its own execution and retrieve the specific runtime state needed to diagnose it. This points toward execution history being used not only to explain agent behaviour after the fact, but potentially to help the agent identify where its own workflow went wrong and what should be revisited.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.