acceptodds
Under review as a conference paper at ICLR 2027

AEGIS: Adversarial exploration for generalized improved state representations

Abstract

Reward-free pretraining (RFPT) seeks transferable state representations and exploration behaviors before task-specific rewards are available, but existing approaches treat exploration and representation learning as independent processes. AEGIS combines them as an approximate bilevel Stackelberg problem between (i) a latent dynamics model that minimizes the prediction error and (ii) an exploration policy that uncovers states that expose its weaknesses. The policy is driven by an intrinsic reward combining intra-episode diversity with novelty toward states the dynamics model has not yet fitted. We prove that each component of the representation model and the intrinsic reward exclude a degenerate solution the others allow, so only their combination is non-degenerate. AEGIS outperforms state-of-the-art methods with faster convergence and less redundant exploration on MiniGrid and Procgen.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.