acceptodds
Under review as a conference paper at ICLR 2027

Suppressing a Rejected Action Costs a KL: Closed-Form Test-Time Updates for LLM Agents

Abstract

Large language models are widely deployed as agents that interact with external environments through structured actions such as shell commands, tool calls, or search queries. Once the model emits an action that the environment cannot execute, the interaction step is interrupted. In-place recovery then faces a dilemma: with frozen weights, stateless resampling tends to fall back into similar invalid outputs and repeat the same failure; with weight updates after every rejection, gradient-based adaptation introduces substantial latency and compute cost. We propose TALE (Test-time Adaptation via Likelihood Enforcement). TALE turns an observed rejection into an exact likelihood-suppression target and applies a low-rank representation edit at the model's output side; the update direction is determined analytically and the suppression intensity is fixed by a one-dimensional monotone root-finding procedure, with no backpropagation. The edit reshapes the local probability space and moves subsequent attempts away from the failed neighborhood. Across four benchmarks covering terminal commands, tool calls, and web search, at three model scales, TALE reduces the unresolved rate of interaction steps relative to update-free resampling and saves a large fraction of model calls; relative to gradient-based test-time adaptation methods, it attains comparable or higher task success while cutting end-to-end latency by a factor of several, with negligible additional GPU memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.