acceptodds
Under review as a conference paper at ICLR 2027

Asymmetric On-Policy Distillation Learns In-Context Exploration

Abstract

Sequential decision-making under partial observability is difficult because agents must explore to infer hidden, task-relevant information before acting effectively. Learning such information-gathering behaviors tabula rasa via online reinforcement learning is sample-inefficient, whereas offline imitation fails to learn exploratory behaviors due to insufficient coverage of deployment-time histories. We focus on the setting of asymmetric student-teacher learning in partially observed settings. However, direct imitation of privileged experts fails because such experts never need to explore. In this work, we propose ASTEROID, a simple yet surprisingly effective online framework for learning test-time exploration. Our key idea is that history-conditioned student policies can learn to explore through online asymmetric distillation: the student first collects on-policy context under partial observability, and the privileged expert provides action labels conditioned on that context. Iteratively, as this context grows, the student learns to use past observations to infer hidden task information and act according to the inferred latent state, yielding exploration similar to efficient Bayesian posterior sampling. Across a diverse range of eleven different simulated environments, including discrete exploration benchmarks, robotic manipulation, navigation, and Procgen environments, ASTEROID achieves 2 higher returns under fixed exploration budgets and uses up to 100 fewer interactions than baselines. Finally, our method transfers policies trained entirely in simulation to vision-denied real-world robot manipulation, exhibiting efficient exploration and improving success by over baselines. Overall, it provides a simple yet surprisingly effective framework to learn efficient test-time exploration from privileged supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.