From Failure to Foresight: Temporal Anticipatory Reflective Optimization for Skill Learning in Agents
Abstract
abstract Large language model (LLM) agents increasingly rely on reusable skills to solve complex long-horizon tasks. Early approaches use static skill libraries, while recent methods improve adaptability through dynamic skill evolution. However, existing methods remain largely reactive, relying on retrospective error correction after failures occur. This is fundamentally limited in complex environments, where early suboptimal skill choices can trigger delayed and cascading downstream failures. To address this, we propose Temporal Anticipatory Reflective Optimization (TARO), a reinforcement learning framework that shifts skill learning from retrospective correction toward anticipatory prevention. Through iterative execution feedback during training, TARO performs implicit temporal process modeling, enabling a trainable skill generator to anticipate future execution dynamics and proactively avoid downstream dead ends without explicit world models. A retrospective credit assignment mechanism further flattens temporally refined skill variants into counterfactual optimization groups for efficient GRPO updates. Experiments on WebShop and ALFWorld show that TARO consistently improves task success and skill robustness over strong baselines, establishing anticipatory reflective optimization as a promising direction for scalable autonomous skill learning in LLM agents. abstract
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.