acceptodds
Under review as a conference paper at ICLR 2027

When Generalization Goes the Wrong Way: The Limits of Zero-Shot Offline Meta-RL

Abstract

Context-based offline meta-reinforcement learning adapts to new tasks by inferring a latent variable while keeping the policy fixed. Can failures beyond the training task distribution be resolved by better task inference, or do they reflect missing behaviors in the policy itself? Across three continuous-control task families and five learners we observe a pronounced asymmetry: magnitude extrapolation often retains useful performance, whereas direction extrapolation can produce returns below a stationary baseline. We analyze this failure by separating latent-selection error from limitations of the frozen policy's behavioral repertoire. In Ant-Dir, the measured heading range over evaluation tasks matches that attained on training tasks in of models; broadening behavior-data coverage at fixed training tasks yields little expansion; and return-based latent search fails to recover three distant goals while recovering two in a full-support control. Together these results support a behavioral-coverage bottleneck that task inference alone does not overcome in any condition we test. Strong in-support returns can meanwhile mask severe behavioral collapse, so we recommend direct behavioral probes alongside return-based evaluation of zero-shot offline meta-RL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.