Implicit Strategic Optimization: Prediction-Aware Long-Horizon Learning for LLM Agents
Abstract
Learning effective strategies in imperfect-information games requires agents to account for delayed consequences while adapting to partially observed opponent behavior. Although such interactions appear non-stationary, they often contain recurring strategic regimes that can provide useful structure for learning. We introduce Implicit Strategic Optimization (ISO), a prediction-aware framework for long-horizon learning in large language model (LLM) agents. ISO combines context prediction, delayed-value estimation, and routed policy optimization. Its central mechanism separates the context used for acting from the context used for updating: a predictor guides action selection from pre-action interaction history, while context labels logged after decisions determine which context-conditioned policy components receive updates. A Strategic Reward Model estimates delayed returns from trajectory segments to guide routed group relative policy optimization. We analyze this routing principle in an idealized latent-context repeated game, deriving a contextual regret bound that separates context-specific initialization, prediction errors, and within-context variation. Experiments in No-Limit Texas Hold’em and competitive Pokémon demonstrate improved strategic performance over supervised fine-tuning and reinforcement learning baselines. Ablations support the contributions of both delayed-value estimation and context routing, while evaluations on unseen and hybrid poker opponents show that the return advantage extends beyond the training population. Behavioral analyses further associate the gains with decisions that trade immediate payoff for future strategic value. These findings support prediction-aware context routing as a useful mechanism for long-horizon learning in strategic LLM agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.