acceptodds
Under review as a conference paper at ICLR 2027

REINFORCEMENT LEARNING FROM ENVIRONMENT KNOWLEDGE

Abstract

Reinforcement learning with verifiable rewards (RLVR) has proven effective in training large language models (LLMs) for long-horizon agentic tasks. In specialized domains, however, agents often need to acquire unfamiliar environment knowledge to achieve expert-level performance. For domains where expert demonstrations are scarce but unstructured materials such as documentation are present, RL enables the model to explore on its own and acquire new capabilities through environment interaction and feedback. Yet existing RL recipes mainly optimize for the solution's score, which discourages knowledge acquisition when early attempts to apply unfamiliar techniques lower it. To bridge this gap, we introduce Reinforcement Learning from Environment Knowledge (RLEK), an RL-based method that encourages knowledge acquisition alongside optimizing the solution score. Concretely, we augment vanilla performance episodes with dedicated exploration episodes, each of which starts from a solution the model itself produced earlier in training and is rewarded for successfully adopting a missing documented technique. We study RLEK in GPU kernel generation for a new hardware generation whose instruction set is largely unfamiliar to the model, while relevant documentation is available in the environment. On six held-out GPU kernel tasks, the RLEK-trained model uses documented new techniques in 12% of its passing kernels, compared with none under vanilla RL. RLEK achieves higher best-of-16 performance on five tasks and matches vanilla RL on the sixth, while maintaining similar median performance among correct solutions on most tasks. On Kimi Delta Attention, it achieves the highest median and best-of-16 performance among the evaluated models, exceeding Claude Opus 4.8 under matched context length and time budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.