Zero-Shot Control with Gaussian Representations and Test-Time Adjustment
Abstract
Zero-Shot Reinforcement Learning (ZSRL) pretrains general-purpose policies from reward-free experience to enable deployment on unseen rewards without further learning. Successor Feature (SF)-based methods facilitate this transfer by fitting a linear reward coefficient to retrieve a pretrained policy. However, existing approaches often overlook how representation geometry affects SF evaluation, while direct policy retrieval can be suboptimal under reward misspecification and function-approximation errors. First, we introduce Gaussian Latent Dynamics-based Successor Features (GLDSF), which combines latent dynamics prediction with regularization toward an isotropic Gaussian to learn control-relevant representations with well-conditioned geometry for SF evaluation. To mitigate retrieval suboptimality, we propose a model-based test-time adjustment procedure that searches over the pretrained policy family rather than over action sequences, with no parameter updates or additional environment interaction. We provide analysis supporting this adjustment design. On the ExORL benchmark, planning-free GLDSF is competitive with the strongest state-based baseline and outperforms the strongest pixel baseline by 24.0% in aggregate. Incorporating test-time adjustment further boosts aggregate performance by 7.1% on states and 8.4% on pixels. Ablation studies further examine the effects of representation regularization, search space, and other test-time adjustment design choices.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.