acceptodds
Under review as a conference paper at ICLR 2027

ORBIT: Online Regret-Bounded Intervention for Tool-Using Agents

Abstract

Tool-using language-model agents can cause harm through a sequence of actions, even when no single action looks suspicious. Some recent defenses constrain agents to pre-specified or permitted execution paths. We introduce ORBIT (Online Regret-Bounded Intervention for Tool-Using Agents) that instead learns when to intervene from evidence accumulated across an agent's actions, without requiring a complete execution plan in advance. It fits an allow-inspect-block policy over calibrated signals from several monitors, including two LLM critics from different model families. ORBIT can inspect to gather more evidence before committing, deferring the allow/block choice at ambiguous steps, building evidence memory across connected episodes from beliefs about persistent compromise, current deviation, and progress toward harm. Our experiments show that ORBIT prevents harm under varying levels of evidence restriction and under adaptive attack sequences. In action-branch replay, it stops markedly more attacks early than DBN posterior thresholding at similar benign cost. On fixed AgentDojo Workspace recordings from live OpenAI and Anthropic executions, carrying evidence across episodes improves mean prevention over episode-reset controls, at the cost of interrupting more legitimate work; the advantage is largest when observation is sparse and the taint channel is removed. Under adaptive attacks that revise the payload across successive attempts, persistent memory lowers harmful goal completion relative to resetting memory each episode. An exploratory study also finds that updating beliefs from modeled monitors' replies to intervention recommendations increases mean early stopping, at higher benign cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.