acceptodds
Under review as a conference paper at ICLR 2027

Long-Horizon Insurance Credit: Measuring Preventive Actions in Long-Horizon Tasks

Abstract

AI agents can sometimes make themselves more reliable by taking preventive actions, such as saving a checkpoint or writing down useful information. These actions may look useless when everything goes smoothly, but they matter a great deal if the agent later loses context, forgets earlier work, or encounters a failure. This makes preventive actions hard to evaluate: credit signals evaluated only under clean conditions can miss this value, while simple fault-injection tests can also give a misleading picture. We introduce Long-Horizon Insurance Credit to estimate the value of these actions. It measures how much an action helps when failures occur, then subtracts what the same action changes when nothing goes wrong. To estimate this quantity fairly, we compare matched runs driven by the same sampled fault tapes. In a controlled setting where the true answer is known exactly, insurance credit identifies useful preventive actions accurately (Kendall's τ = 0.92), while standard clean-run methods miss them. We then test it on three tool-use benchmarks: ToolMaze, BFCL, and ALFWorld. We find that a single checkpoint can be extremely valuable, but current LLM agents usually fail to benefit from it because they rarely check the saved workspace after losing context. When the system makes them look once, much of the benefit returns, with insurance credit ranging from 6.4 to 20.8 percentage points relative to a length-matched placebo on ToolMaze and BFCL. This suggests that the main bottleneck is discovery rather than use: the agents can often exploit saved information once they look, but they rarely choose to look on their own. In this regime, system-level recovery design appears more reliable than relying on larger models alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.