What In-Context Reinforcement Learners Infer About Unseen Tasks
Abstract
In-context reinforcement learning agents adapt to new tasks from a prompt alone, and are widely reported to generalize to unseen tasks. We ask where an agent's implicit task estimate goes when the test task lies outside the convex hull of the pretraining task parameters. We build region-disjoint parameter splits over locomotion and manipulation families and decode that estimate three independent ways: linear probes on rollout-time hidden states, inversion of emitted actions against per-task expert policies, and direct behavioral readout. All three agree that the estimate saturates at the hull boundary rather than failing erratically: an agent given a task beyond its training range acts as though it had been given the nearest task inside that range. The confinement is not a data limitation, since model-free audits show that the context identifies the out-of-hull task exactly, and it is not specific to attention, since classical context-encoder agents pin their task codes at the same boundary. Correcting the estimate does not repair the failure and can invert it: supplying the true out-of-hull task code degrades return below that of the boundary-pinned agent, because the policy is competent only at conditioning codes it saw in training. Nor is the boundary an artifact of coverage: retraining on exactly relabeled tasks that extend the hull moves the saturation point to the new boundary and no further. Hull confinement thus bounds what these agents infer, and repairing the inference alone does not restore what they can do.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.