acceptodds
Under review as a conference paper at ICLR 2027

StyleCraft: Latent Augmentations as a Step Toward Generalizable Reinforcement Learning

Abstract

Reinforcement learning agents often perform well in their training environments but degrade under visual distribution shifts. Existing approaches commonly address this problem with predefined pixel augmentations, whose effectiveness depends on how well the selected transformations cover unseen visual variation. We propose StyleCraft, a two stage framework that combines task style disentanglement with learned visual augmentation. During the first stage, pixel augmentations together with gradient reversed reconstruction encourage the policy visible task representation to discard appearance specific information while a separate style representation captures visual variation. Once the generative model has learned a sufficiently structured style space, the second stage uses interpolation between style representations to construct new visual observations while keeping the task representation fixed. These generated observations replace fixed pixel augmentations in the invariance objective and provide additional supervision for learning a more visually robust task representation. Experiments across KAGE Bench, DCS based locomotion, RL4VLA, and ManiSkill show that StyleCraft improves robustness to unseen visual conditions across controlled visual shifts, continuous control, and robotic manipulation. We additionally reproduce the ManiSkill scene on a real robot, where StyleCraft retains nonzero success under unseen table colors while the PPO baseline fails.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.