acceptodds
Under review as a conference paper at ICLR 2027

Excitation Collapse in In-Context Control: Symmetric Freezing and Symmetry-Broken Imitation

Abstract

In-context controllers are sequence models pretrained on families of dynamical systems so that they can control an unseen system from its interaction history. They are usually trained to imitate an oracle that knows the system, or to regress the oracle's privileged parameters. It is well known in adaptive control that controllers which hedge against parameter uncertainty can turn off: they stop exciting the system and so stop learning about it. We show that imitation pretraining produces this failure exactly whenever the family has a symmetry that the history cannot resolve before the controller acts; unknown actuator polarity is the standard example. The imitation optimum is equivariant, so from a cold start its closed loop stays on the symmetry-fixed actions for all time, which we call symmetric freezing. Because the frozen controller minimises the imitation loss over all history-measurable controllers, changing the architecture or adding capacity and data does not avoid it, and better optimisation moves the learner closer to it. The theory predicts degenerate latents in privileged latent distillation, a default direction of predictable sign under asymmetric priors, and exact freezing of equivariant architectures. We confirm each prediction for discrete and continuous symmetry groups and introduce two black-box diagnostics. Since the failure comes from the training target, we change the target. Symmetry-broken imitation (SBI) trains an unmodified network on a teacher that commits to one orbit element and revises this choice using exact privileged likelihoods. The teacher is provably consistent, and one round of on-policy relabelling transfers its revisions to the student. With one plain transformer and no knowledge of the symmetry at deployment, SBI removes all instability in a linear family (two seeds, 180 paired runs each), matching an explicit-copy quotient controller. In RMA-style latent distillation on a nonlinear pendulum with unknown motor polarity, changing only the regression target and adding one round of on-policy relabelling reduces falls from 32–43% to 10–13% across two seeds (p < 10^-12 in each), on par with classical multiple-model control. In a six-state family whose prior no model bank can cover, SBI beats multiple-model control with 1024 models on paired runs across three training seeds (p = 8 × 10^-12 pooled).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.