acceptodds
Under review as a conference paper at ICLR 2027

Learning Task Representations from Sparse Evidence in Offline Meta-Reinforcement Learning

Abstract

Context-based offline meta-reinforcement learning (COMRL) trains a context encoder to infer a latent task from collected trajectories. Existing methods commonly treat task-relevant evidence as broadly distributed across the context, or isolate a few individually informative transitions. This leaves open the case in which only a small fraction of interactions carries task-dependent information, while the rest is valid but weakly informative experience. A trajectory may still contain sufficient task information, but the informative feedback can appear in brief intervals of unknown location and duration. Learning a reliable task representation then requires combining these sparse clues while limiting interference from the rest of the context. We study this sparse-evidence regime and show that task inference requires preserving weak, collectively informative responses before determining their relevance. Based on this observation, we introduce STeR (Slot-Based Task Representation Learning from Sparse Evidence in Offline Meta-Reinforcement Learning), a structured context encoder that performs accumulation before selection. Transition features are softly assigned to content-based latent slots, each slot is normalized independently, and a task-level readout selects the summaries that matter. The organization is learned from trajectory-level task labels alone, with no evidence masks or temporal localization supervision. We evaluate STeR on controlled offline meta-reinforcement learning environments where task information appears through sparse rewards, transient motion responses, and hidden contact dynamics. Across variations in evidence duration and onset, STeR improves task identification and downstream adaptation while preserving performance when task evidence becomes sparse. The results suggest that structured accumulation is a useful inductive bias for task representations drawn from incomplete, noisy offline experience.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.