Task-Identifiability-Aware Action Supervision for Offline In-Context RL
Abstract
In-context reinforcement learning (ICRL) enables policies to adapt to unseen tasks through accumulated interaction context without test-time parameter updates. However, existing approaches typically apply uniform action supervision across context prefixes, overlooking differences in how reliably each prefix identifies the underlying task. A short or ambiguous prefix may remain consistent with multiple tasks that require different actions for the same query state, such that the target action cannot always be uniquely inferred from the available context. To address this mismatch, we propose Tiara (Task-Identifiability-Aware Representation and Action Learning), an offline ICRL framework that makes task inference explicit and uses empirical task-identification reliability to guide action learning. Tiara employs a causal context encoder, independent of the query state, to construct prefix-wise task representations supervised by training-task identities. The correctness of auxiliary task identification serves as an empirical proxy for prefix-level task identifiability and is used to reweight action supervision, assigning greater weight to correctly identified prefixes while retaining a smaller positive weight for ambiguous ones. Experiments across diverse multi-task offline ICRL benchmarks show that Tiara improves generalization to unseen tasks over strong baselines across most settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.