acceptodds
Under review as a conference paper at ICLR 2027

Offline Learning to Diagnose and Intervene in Tool-Using Agents

Abstract

Tool-using agents act on stateful environments, where erroneous tool actions, such as API calls or file operations, can leave persistent effects that are not directly visible in later interactions. This creates a sequential safety problem: an external harness must track possible hidden damage, decide when costly diagnosis is worthwhile, and choose how to intervene while accounting for intervention cost. We study how to learn such a harness from fixed interaction logs in which damage labels are revealed only selectively. We propose Revealing-Action Pessimistic Value Iteration (RA-PEVI), which learns intervention dynamics shared across diagnostic branches from selectively revealed latent states and performs pessimistic planning in the belief space. We establish a finite-sample performance guarantee, derive the corresponding rate, and prove matching lower bounds and identifiability limits. Experiments on synthetic data and AppWorld support our theoretical results and show that RA-PEVI achieves better performance than standard baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.