When Should an Agent Commit? Preference-Conditioned Workflow Control
Abstract
Stateful LLM agents repeatedly face a choice between replanning and committing to a multi-step plan. Frequent replanning preserves responsiveness to new observations but incurs repeated reasoning cost, whereas longer commitments amortize planning overhead while delaying high-level plan revision. We formulate this trade-off as preference-conditioned adaptive commitment and introduce , a high-level controller that selects the next commitment horizon from the current workflow state and a quality–cost preference while keeping the executor frozen. restores intermediate execution states to compare alternative horizons under a shared frozen continuation, yielding matched downstream quality–cost profiles. Horizon-specific predictors model relative quality retention and computational saving; the requested preference combines these predictions into a commitment policy, while a trainable residual is refined through on-policy interaction. We further analyze the induced semi-Markov control problem and use a finite controlled study to characterize when feedback-dependent horizon selection has value. The benchmark protocol evaluates the learned controller on ALFWorld and WebShop against fixed, query-level, and learned commitment strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.