acceptodds
Under review as a conference paper at ICLR 2027

First Visits and Revisits: Learning from the Deployment Stream in Model-Based Control

Abstract

A controller facing recurring dynamics meets first visits to new dynamics and revisits to dynamics that its stream has shown. We argue that adaptation then runs on two time scales, fast selection among hypotheses the model holds and slow learning of those hypotheses, which compete for the same weights. We propose FSSL (Fast Selection, Slow Learning), a model-based controller whose members keep a frozen copy of their pretrained prediction heads beside a copy that one multiple-choice loss trains between episodes. Each member plans with the head of lowest recent error among both copies. On six streams against eight baselines, FSSL gives the best mean on all six, and slow learning, hard credit and per-head weighting each raise its score significantly. Without the frozen copy, the learning heads lose on first-visit entries and gain on revisits of CartPole and HalfCheetah, whereas with it a frozen head best explains the entered mode for 79 to 100% of the members. Planning-error bounds per phase explain this division of labour. Code and demo are available at anonymous.4open.science/r/fssl.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.