acceptodds
Under review as a conference paper at ICLR 2027

JEPath: a world model for first-person navigation

Abstract

JEPA world models combined with test-time trajectory optimization have recently achieved strong results on tasks conditioned on a single goal image. However, most of these results assume fully observable environments, and their adaptation to first-person navigation raises specific challenges. In this study, we introduce PADoom, a benchmark built on ViZDoom where the same environments and trajectories can be observed either from a fully observable top-down view or from a first-person view. We show that perceptual aliasing, where different locations produce nearly identical observations, requires combining visual and spatial information, and that the spatial signal must avoid absolute position encoding to generalize across maps. Based on these findings, we propose JEPath, a new JEPA world model that jointly predicts future visual features and ego-motion from observations and actions. Unlike models that predict absolute position, JEPath predicts motion and integrates it along candidate paths. The planner scores paths by distance to the goal rather than visual similarity, resolving perceptual aliasing while still relying on vision to handle navigation-relevant scene structure such as obstacles. On PADoom, JEPath outperforms recent world models and goal-conditioned policies for first-person navigation. It also generalizes substantially better to unseen maps, where competing methods experience a pronounced performance drop. The benchmark and models will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.