Offline Meta Reinforcement Learning with Bayesian Task Relabelling
Abstract
Offline Meta-reinforcement learning faces two challenges when transferring from collected data to online adaptation. Behaviour-dependent task inference arises when an encoder identifies tasks from collector-specific patterns rather than task properties. Belief distribution shift can lead to poor actions when the policy encounters unfamiliar beliefs about the task. We introduce Bayesian Task Relabelling (BTR) to address both challenges using a learned Bayesian task model. We build on the Bayesian linear task model and identify how its conditional reward and next-state likelihoods exclude collector behaviour as separate evidence of task identity. To broaden belief coverage, BTR samples imaginary tasks using a posterior-based approximation and generates their corresponding imaginary rewards along collected histories. The same model combines these imaginary rewards with collected next states to update task beliefs in closed form, providing additional data for policy training. Task sampling is not restricted to training tasks or achieved goals, and reward generation requires no known task-specific reward function. In comparisons with recent or representative baselines trained on the same data, BTR achieves the highest mean return on all eight control benchmarks under online adaptation and seven under offline-context evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.