acceptodds
Under review as a conference paper at ICLR 2027

Embodied Learning in the Loop: Co-Adapting Policies and Environment Distributions

Abstract

We devise a learner-aware framework for embodied PointNav that treats the scene distribution as a policy-conditioned component of training. Existing policies are trained on static corpora or samples from pretrained generators, neither of which responds to the learner’s evolving competence. Moreover, scene-centric properties (e.g., realism, reachability, geometric complexity) do not identify scenes that promote learning. Our framework adapts a layout diffusion generator using scene-level policy learning progress. It generates indoor layouts for PointNav and compares policy performance on the same episodes before and after each policy update. Changes in success, path efficiency, and goal distance define a utility signal that weights the denoising records that produced each scene, steering updates to the pretrained generator. This closes the loop between scene generation and policy learning. Averaged across five distributions, our pipeline improves PointNav performance (5.84 and 4.44 percentage-point gains in SR and SPL, respectively, over Frozen Generator). Policy feedback shifts the generator to a distinct distribution (MMD²=0.752), with the separation persisting across later adaptation intervals. Matched-success trajectories and same-input comparisons further reveal distinct navigation behaviors and action choices under identical observation and action histories. Our code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.