Source-MSIS: Tracing Source Errors in Interactive Agents
Abstract
Accurately tracing root causes is key to helping interactive agents recover from failures in long-horizon tasks. Observations at failure may reveal only the cascading effects of earlier, unnoticed errors. Existing methods can flag suspicious steps in execution traces, but struggle to verify root causes and derive precise, nonredundant recovery plans. We propose the Source-indexed Minimal Sufficient Intervention Sequence (Source-MSIS) framework for tracing source errors and verifying recovery. The framework uses execution records from failed tasks to identify earlier errors that may block later steps and proposes targeted repairs for these errors. It then replays the task to check whether the repairs remove the corresponding blockages. If the task remains incomplete, feedback from the new execution guides the search for other errors and additional repairs. Once the task is completed, an independent verifier removes one or more repairs and replays the remaining repairs in their original order under the same conditions. A plan passes verification only if the full plan completes the task and every reduced plan fails to do so. Across three interactive benchmarks with publicly available base models, Source-MSIS achieves an overall recovery success rate of 65.3%, compared with 5.7% for the strongest baseline, and reaches 100% on ALFWorld. These results show that tracing source errors and verifying recovery through replay provide concrete targets for repair, helping agents complete previously failed tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.