acceptodds
Under review as a conference paper at ICLR 2027

A Fix Is Not a Rule: Learning When to Reuse Agent Corrections

Abstract

Large language model agents can learn from failures, generate behavioral corrections, and retain them for future tasks. Yet successfully fixing one failure does not determine when a correction should be reused: the same correction may help, have no effect, or cause harm across different tasks. Agents therefore need to learn both how to correct mistakes and the conditions under which a correction is useful. We propose an experience selection method based on correction scope learning. By observing the direct effects of the same correction across task contexts, the method predicts its utility and guides selective reuse. Experiments on LifelongAgentBench and WebShop, together with supplementary experiments using GLM-5, DeepSeek-V4-Flash, and Qwen3.6-27B, show that the same correction can have positive or negative effects across tasks. Further experiments with GLM-5 show that, under the same activation budget, HCAR-Risk improves mean deployment gain by 0.642 per correction over activation-matched random selection and increases the proportion of selected correction–task pairs with positive effects from 51.9% to 87.8%. These findings establish correction scope as a key learning problem in agent experience reuse: agents need to accumulate effective corrections and use direct evidence of their effects to decide when to apply them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.