acceptodds
Under review as a conference paper at ICLR 2027

Enhancing Large Language Models with Decision-Relevant Knowledge for Reinforcement Learning Tasks

Abstract

Large language models (LLMs) have demonstrated strong language understanding and reasoning capabilities through large-scale pretraining, motivating their increasing use in reinforcement learning and sequential decision-making. However, pretrained LLMs lack sufficient understanding of task-specific decision-relevant knowledge, including critical state information and task rules that directly guide action selection. While such knowledge is essential for effective decision-making, general-purpose LLMs often fail to correctly interpret and leverage it in domain-specific reinforcement learning tasks. To address this gap, we propose a two-stage training framework that explicitly teaches LLMs task-specific decision-relevant knowledge before policy optimization. In Stage-A, we construct supervision signals from structured-state definitions and task rules to train the LLM to recognize key environmental facts and infer the current decision objective, yielding a decision representation . In Stage-B, a lightweight policy is optimized with PPO on top of through online interaction. Under the same online interaction budget, our method achieves better overall performance than the compared baselines on Atari and Crafter and also shows applicability to additional environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.