acceptodds
Under review as a conference paper at ICLR 2027

Learning to Allocate: Critic Training for Test-Time Agent Search

Abstract

Test-time search for LLM agents requires deciding which partial trajectories deserve further computation. Standard critic training supervises each prefix independently using continuation outcomes, although the resulting scores jointly determine how a shared search budget is allocated. We introduce RATE (Relative Allocation Training), a framework that combines pointwise continuation-value learning with relative supervision across competing prefixes. Both forms of supervision use the same sampled continuations. The joint objective preserves the ideal continuation-potential target while emphasizing relative score errors that affect allocation. At test time, the trained critic guides the allocation of continuation slots while the actor remains fixed. Across four agent environments, RATE improves mean candidate reward over Best-of-N (BoN) by 3.95–8.90 percentage points under a fixed continuation budget. The gains also extend from fixed-frontier evaluation to repeated budget allocation during online search.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.