acceptodds
Under review as a conference paper at ICLR 2027

WorldPO: Holistic Agentic Learning with Policy Optimization and World Modeling

Abstract

Although pretrained large language models possess broad capabilities, limited environment-specific knowledge can hinder their adaptation to unfamiliar agentic environments. Policy optimization improves behavior through task feedback, yet uses information-rich environmental observations primarily as decision context, leaving their value for environmental understanding underexploited. We propose **WorldPO** for *holistic agentic learning* that integrates behavioral improvement with environmental understanding through collaborative policy optimization and world modeling. Using the same on-policy rollouts, WorldPO combines feedback-guided policy learning on generated responses with world modeling supervised by environmental observations, optimizing both objectives in a shared LLM. We evaluate WorldPO in three environments with distinct observations: ALFWorld, WebShop, and Search. WorldPO improves task completion and complements GiGPO's action-level optimization, raising 3B ALFWorld success from 75.00% to 92.19%. On Search, 3B WorldPO outperforms GiGPO on all four non-training datasets. Our analyses suggest that WorldPO benefits from task-relevant environmental knowledge beyond exact action-observation correspondences, with greater gains from successful experience. WorldPO also reduces invalid actions in ALFWorld and uses fewer turns on jointly solved tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.