acceptodds
Under review as a conference paper at ICLR 2027

Group-Context Symmetry Breaking: Turning Environmental Asymmetries into a Learnable Group Context for Reinforcement Learning

Abstract

Symmetry lets reinforcement learning agents share experience across related states and actions, but environmental asymmetries make that sharing biased. How can a broken symmetry remain a useful structural prior? We propose Group-Context Symmetry Breaking (GSB), which represents the asymmetry in a learnable context that transforms with the task's group. GSB enforces joint equivariance over state, action, and context, so an asymmetric environment becomes a fixed-context slice of an equivariant family. Real transitions teach the model which slice to use. We derive bounds linking context displacement under the group, model error, and symmetry deviation. Any model equivariant in state and action alone must pay at least half the largest deviation, so GSB provably approximates the environment better whenever its context family's approximation error falls below that floor. In soft actor–critic, the learned model generates experience in locomotion, and in manipulation it corrects real transitions carried along the group orbit while the policy reads the learned context. We compare GSB with strict, relaxed, and partially equivariant priors, including partially equivariant SAC built for symmetry-breaking environments, and with unstructured SAC. Under controlled symmetry breaking, these priors fall below even unstructured SAC on most tasks, whereas GSB attains the highest mean return of all methods at the end of training on the Hexapod, Ant, and Humanoid tasks and the most consistent return on Walker2d. On two manipulation tasks where agents train on one half of the goal domain under a constant lateral force, GSB attains the best held-out return against every baseline and ranks first in every seed. GSB's correction of the transported transitions at the learned context adds 4.35 and 1.76 in held-out return, each larger than GSB's lead over the strongest baseline, so the lead rests on this correction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.