Where Did the Agent Go Wrong? Localizing—and Fixing—Failure Steps in Multi-Step Video-Editing Agent Traces
Abstract
Autonomous agents run long, tool-rich workflows, yet when a run fails, telling where it went wrong is hard: a syntactically valid trace can quietly drift off the user's goal. We study this failure-step localization problem in multi-step video editing, where one request expands into a chain of metadata probes, shot detections, clip cuts, and concatenations. We introduce AV-Fail, a benchmark that asks a judge to identify the earliest faulty step of an agent trace. AV-Fail couples a controlled track injecting one of five fault types (or none) for programmatic exact-step gold labels with a real track of genuine runs verified via multi-model human-in-the-loop review. A blind two-annotator study fixes the human ceiling at Cohen's (substantial), which no automatic judge reaches: the best single model attains only and different judges err in opposite directions, so automatic failure localization cannot stand in for human review. Across five judge models we find a sharp difficulty spectrum: mechanical faults are nearly solved (), a cumulative-budget fault is a systematic blind spot (), and clean traces trigger over-attribution (a false-positive rate); both effects are strongly model-dependent. Two lightweight mitigations each help but interfere when combined. We resolve this with Trace-State Prompting, which—wherever a constraint is expressible as a checkable predicate over the trace—computes that trace state deterministically and routes only the residual decision to the model, raising accuracy from to and closing the budget blind spot losslessly. On real traces it surfaces semantic failures that objective signals miss, showing off-goal drift—not crashes—is the dominant unhandled failure mode. The findings are not video-specific: ported to a data-pipeline (ETL) agent domain, the same ordering recurs and the protocol again lifts accuracy from to . AV-Fail provides a reproducible lens on a capability that agentic evaluation has so far left largely unmeasured.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.