Recover or Cascade? The Hidden Role of KL Regularization in Multi-Turn LLM Agents
Abstract
Preventing errors from propagating through a multi-turn language-model agent is important for reliable task execution. However, evaluations of agent training commonly report accuracy and final task success, which do not distinguish avoiding errors from recovering after them. We study how on-policy reinforcement learning (RL) shapes behavior after errors by varying the coefficient in penalties motivated by Kullback–Leibler (KL) regularization. The intended constraint on departures from a reference policy provides a way to examine whether training preserves existing behavior or allows corrective actions to be learned. We measure failure persistence, the probability of failing after a failed step, and sequential consistency, the probability of passing after a passed step. Across five main benchmark settings, failure persistence is higher at than at in four. Our analyses further reveal that, as changes, higher sequential consistency can coexist with a lower probability of returning to a passing step after failure, distinguishing gains in continued correctness from gains in recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.