Self-Calibrating Interruption of Agent Doom Loops via Risk-Controlled E-processes
Abstract
Language-model agents run in harness loops that call an API-served policy and execute tools, so every iteration is billed. Much of that cost buys nothing: agents fall into *doom loops*, spending long after success has become improbable. Turn caps, token budgets and repetition heuristics kill salvageable episodes, per-step judges add a model call per step, and activation probes need internals that an API does not expose. The target is undefined: the conditional success probability is a zero-drift martingale, so no rule can wait for it to decline. Because a controller inspects every step, fixed-sample calibration is either invalid, realizing a false-kill rate at a nominal , or crippled by a horizon union bound that cuts cost by only . Least recognized, an interruption policy censors its own labels: the episodes it kills are exactly those whose uninterrupted outcomes identify its risk. We present Circe, which defines a doom loop as a *value-stall* (low salvage probability with vanishing information yield per dollar) and certifies it with a test supermartingale over a log-derived stall indicator, bounding by the probability of ever halting a salvageable trajectory under adaptive inspection, with no horizon penalty and asymptotically optimal delay. Its central mechanism is *randomized reprieve*: a firing certificate halts with probability and otherwise releases the episode to natural termination with known propensity , restoring identification and yielding a doubly robust salvage-loss estimator with a time-uniform confidence sequence and self-calibration regret over deployment rounds. On software-engineering, terminal and web-navigation suites and a near-saturated tool-calling control with `glm-5.2`, Circe cuts cost at a pp pass@1 difference and false kills, using log-derived features and no model internals, extra calls, or GPUs. Under a mid-stream model and difficulty shift every tested baseline breaches its risk constraint by up to unnoticed, while scheduled reprieve caps violation at and flags it within two rounds at exploration spend.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.