acceptodds
Under review as a conference paper at ICLR 2027

Reuse, Refresh, or Replan? Selective Decision Updating for Real-Time LLM Agents

Abstract

Real-time large language models (LLMs) agents must decide whether to reuse a previous decision or update it as the environment evolves. Reuse saves computational cost when the decision remains valid, but an inherited action can mislead the agent once the state changes. Controlled intervention experiments show that prior-action carryover benefits agent performance under safe conditions, yet induces decision anchoring once the action becomes unsafe. Explicit action-outcome information improves the selection of alternative actions. Motivated by these observations, we study selective decision updating supervised via simulator replay. A low-cost model predicts the action for the current state alongside an action-treatment label, while a controller estimates whether further reasoning could correct the action before fallback. The controller triggers AgileThinker if the candidate is unusable or its estimated risk exceeds a threshold selected through validation. On held-out seeds and states, selective fallback trails full AgileThinker by 0.6 percentage points in one-step reference-set accuracy, while reducing mean completion tokens by nearly 60% and invoking the thinking model on about 44% of requests. Collectively, these results show how the expected benefit of further reasoning can guide computational resource allocation for agents updating decisions in dynamic environments, within the evaluated offline, one-step task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.