Exploratory Prompting for In-Context Reinforcement Learning
Abstract
In-context reinforcement learning pretrains a sequence model on many tasks' interaction data and deploys it frozen, expecting it to improve on a new task from the episodes it adds to its context. We show that a cold start can trap the agent in that context. Acting with its mean action, the model plays a default episode from an empty context. The context it then writes can repeat or amplify that episode, although the episode's rewards still carry the task. We evaluate seven in-context RL methods on ten task families and dissect the trap in the family where a stuck trial is unambiguous. Context swaps in frozen models show that the model reads the task from the rewards of an episode it did not write. A random episode releases almost every stuck trial, and relabeling its rewards steers the agent toward the relabeled task. Our remedy, exploratory prompting, needs neither retraining nor a demonstration: the agent spends one episode acting at random. When its context is a single prompt episode, it keeps that episode as the prompt only if the return it induces is at least the return its own first episode induces. For four methods, exploratory prompting removes most stuck trials and raises the mean final return in 34 of 40 (family, method) pairs by mean gain, at some cost in cumulative return. Separately, for single-prompt methods, dropping the prompt at random during a short fine-tuning stage also removes most stuck trials, whereas step-matched continued training does not. We argue that part of the reported difference between in-context RL methods reflects how they start, not how well they learn.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.