acceptodds
Under review as a conference paper at ICLR 2027

SteerControl: World Model-Guided Adaptation of Frozen Policies to Changing Dynamics

Abstract

Unexpected changes in dynamics are often unavoidable for intelligent agents operating in real-world settings. Policies trained under fixed dynamics can suffer substantial performance losses after deployment when the same action produces a different state transition. Many existing adaptation methods address such failures by training on anticipated variations or by relying on deployment rewards, reference trajectories, or parameter updates. We address this problem using a learned dynamics model that requires no supervision beyond what is available in a standard reinforcement learning (RL) setting. SteerControl pairs a frozen policy with a frozen one-step dynamics model, both trained only under nominal dynamics. It uses the difference between predicted and observed transitions to estimate a bounded correction to the policy’s action mean without requiring fault labels, external supervision, deployment rewards, or weight updates. Across twelve continuous-control environments with varied actuator faults and physical perturbations, SteerControl recovers return lost by the faulted base policy under constant and switching perturbations. These results suggest that a nominal dynamics model can help an existing model-free policy maintain more of its performance when faced with unseen shifts in action effects during deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.