acceptodds
Under review as a conference paper at ICLR 2027

When Should an Agent Commit? Preference-Conditioned Workflow Control

Abstract

Stateful LLM agents repeatedly face a choice between replanning and committing to a multi-step plan. Frequent replanning preserves responsiveness to new observations but incurs repeated reasoning cost, whereas longer commitments amortize planning overhead while delaying high-level plan revision. We formulate this trade-off as preference-conditioned adaptive commitment and introduce , a high-level controller that selects the next commitment horizon from the current workflow state and a quality–cost preference while keeping the executor frozen. restores intermediate execution states to compare alternative horizons under a shared frozen continuation, yielding matched downstream quality–cost profiles. Horizon-specific predictors model relative quality retention and computational saving; the requested preference combines these predictions into a commitment policy, while a trainable residual is refined through on-policy interaction. We further analyze the induced semi-Markov control problem and use a finite controlled study to characterize when feedback-dependent horizon selection has value. The benchmark protocol evaluates the learned controller on ALFWorld and WebShop against fixed, query-level, and learned commitment strategies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.