acceptodds
Under review as a conference paper at ICLR 2027

Latent Critic, Pixel Policy: Goal-Conditioned Control That Reuses a Frozen World Model Only in Training

Abstract

In offline goal-conditioned reinforcement learning, a learned value weights the imitation of dataset actions, so it must tell which of two nearby frames lies closer to a goal. Pretrained latent world models are released as frozen checkpoints, which current control methods query at test time to plan over predicted futures. We propose LCPP, which instead uses the frozen world-model encoder in a critic, a value on its latents from which advantage-weighted regression extracts a pixel policy that makes no world-model call. An analysis bounds the reliability of advantage weights by value errors at a state and its successor, including a floor from a frozen input map. At 100k updates on five environments, we compare LCPP with seven from-scratch methods and five methods that reuse the same frozen world model. LCPP is best or tied best in all five columns. In place of the latent value, a pixel value lowers Cube and Scene significantly across 20 seeds per arm, while a latent actor lowers them in each of five seeds. Code and demo are available at https://anonymous.4open.science/r/lcpp.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.