acceptodds
Under review as a conference paper at ICLR 2027

Co-Evolving Acting and Forecasting for Language Agents

Abstract

Training language agents is bottlenecked by the high cost of long-horizon rollouts, yet standard agentic reinforcement learning (RL) learns only from a terminal scalar reward, leaving the supervision in intermediate steps largely unused. Prior work exploits this supervision by predicting next observations, either with a separate world model or as an auxiliary loss on the policy. These objectives, however, are ill-suited to agentic settings: observations such as HTML pages and terminal logs are long and noisy, and single-step prediction mismatches the long-horizon RL objective. We propose CoPE, a simple and effective method that adds multi-step action forecasting as an auxiliary loss. On successful on-policy trajectories, CoPE trains the policy backbone on an action plan containing the current action and K future actions predicted from the earlier history, without the intervening observations. Working in the action space sidesteps environmental noise, and forecasting multiple steps ahead better matches the horizon of policy optimization. We show that this loss optimizes success-conditioned action marginals, so acting and forecasting can be learned jointly within a single shared backbone. Across five benchmarks—ALFWorld, ScienceWorld, WebShop, τ²-bench, and AppWorld—CoPE outperforms GRPO and observation-prediction baselines on aggregate metrics in simulated environments and improves the overall metrics on two realistic applications despite a negative Airline result. It is also up to 1.9× and 2.0× more sample-efficient than ECHO and GRPO, respectively, and requires neither a separate model nor a pretraining stage. A separated-forecaster analysis further supports implicit temporal ensembling through the shared backbone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.