acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Dynamics-Aware Experience Retrieval for Policy Transfer

Abstract

Changes in actuator strength, mass, or friction can make similar states require different controllers. We propose counterfactual dynamics-aware experience retrieval, which compares queries and stored experiences through their responses to shared action sequences. Reference-relative fingerprints capture response direction and magnitude; averaging their distances over each policy's memories selects a controller from a fixed library. Across four DeepMind Control tasks with six policy-training seeds per task, simulator-generated fingerprints reduce mean policy-selection regret by 49-81% relative to random selection and outperform both tested state-similarity baselines in all 24 runs. Same-probe system identification and utility regression remain competitive, with different methods leading on different tasks. With learned visual world models, validation retrieval regret selects models that improve on random selection on Hopper and Walker, whereas held-out rollout loss selects worse-than-random retrievers. These comparisons use separate query draws from the same dynamics grid. The results motivate evaluating world models through the policy choices they support.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.