acceptodds
Under review as a conference paper at ICLR 2027

Playing with Fire: What Transfers When RL Trains a Language Agent?

Abstract

Modern language agents learn specialized skills like safety, mathematical reasoning, and coding with reinforcement learning (RL) and, at inference time, are deployed in diverse, complex harnesses often unseen during training. We study whether RL transfers these learned skills across diverse inference-time settings using two environments: Hanabi, a cooperative hidden-information card game, and K, a concise array-processing programming language, with scaffolds that vary the task knowledge and reasoning support supplied to the agent. We observe that RL gains are spiky: large under scaffolds that supply the task knowledge and small where it is removed. This recurs across model sizes, families, and environments. Supplying the task knowledge after RL helps at least as much as before (+1.75 before and +2.55 after in Hanabi), suggesting that RL teaches the agent to make better use of the knowledge supplied at inference, and not the missing knowledge itself. To investigate what transfers from RL on a hard, verifiable task, we construct FIREWORKS, a dataset of 10K Hanabi belief-state transitions with deterministic verification. Training QWEN3-4B-INSTRUCT-2507 on FIREWORKS elicits longer reasoning traces and improves performance on tasks outside the training objective: supported Hanabi gameplay (+4.3 points) and non-Hanabi reasoning (Pass@8 0.15 → 0.37). Although the resulting checkpoint shows no zero-shot improvement on downstream state tracking (0.117 against 0.127 for the base model), it is a better initialization for downstream RL: on Hanabi it reaches the same state-tracking reward in about 50 optimizer steps against 130 from the base model. RL thus improves how agents use the knowledge available to them and how quickly they learn next, extending its benefits beyond the training task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.