Trust, but Verify: Test-Time Task Inference in Offline Meta-Reinforcement Learning
Abstract
Context-based offline meta-reinforcement learning adapts from a short history of transitions: an encoder turns the history into a task code, and a policy acts conditioned on that code. Outside the training range, return can fall for several reasons: the policy may be unable to solve the new task, the inferred code may be wrong, or the task may lie too far from anything seen in training. We derive an exact four-term decomposition of this lost return and show how to estimate its terms from checkpoint rollouts and data-collecting experts. The measurements reveal two failure modes. On manipulation tasks the frozen policy already solves the new task when given a suitable training code, so the choice of code fails; on locomotion tasks with shifted rewards the oracle-context code is already near the best available and leaves little code-selection headroom, while self-collected context can lose further return. Either way, choosing a code at test time is unreliable: the encoder can infer one from exploration episodes, and a decoder can select the training code that best explains them, but each wins where the other loses and no pre-deployment statistic reliably identifies which one to trust. We propose VeriCode, which explores with a spread of training codes and immediately keeps one seen to succeed. If none succeeds, it compares an encoder candidate and a likelihood candidate through one-episode verification, except when both heads identify the same training code; without a success signal, it compares candidates by return. Under independent evaluation episodes, the probability of selecting the worse of two fixed candidates decays exponentially with the verification budget. Across six Meta-World and six MuJoCo families and three algorithms, VeriCode is at or above the strongest comparator in every cell, raising Meta-World success from 0.52 to 0.83 and MuJoCo return from 200 to 267, at mean adaptation budgets of 5.5 and 9.4 episodes, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.