BRIDGE: Robustness as Dynamic Generalization without Parameter Updates
Abstract
Robust reinforcement learning (RL) uses max-min optimization to withstand worst-case disturbances, often sacrificing nominal performance even when the encountered environment admits a less conservative policy. We introduce *BRIDGE* (Bridging Robustness and In-context Decision GEneralization), which provides an alternative perspective on robustness as dynamic generalization to infer the disturbance from context and *activate a tailored policy without parameter updates*. In-context RL (ICRL) provides a natural implementation, but requires pretraining experience that teaches how disturbances affect the environment and how to counteract them. Our key insight is that the max-min training of canonical robust RL already generates this experience through a curriculum of adversaries and specialized intermediate policies. Motivated by this, we propose the *CORE* pretraining pipeline to distill these experts into an ICRL model and use the final robust controller to collect its initial deployment context. These initial interactions preserve the robust baseline's behavior with minimal risk while providing context for tailored control. Having used ICRL to reduce conservatism around a single nominal task, we take BRIDGE one step further: using adversarial experience to strengthen ICRL's own generalization to unseen tasks under deployment disturbances. This introduces a critical challenge that the context must now reveal both what a new task requires and how its execution is disturbed. To support this joint inference, we propose the *ADAPT* architecture for policy transformers to learn to disentangle task and disturbance information from CORE curricula across training tasks. Our analysis quantifies when contextual specialization improves worst-case return over a conventional robust policy, including on unseen nominal tasks. Experiments on Dark Room, Meta-World, and 11 continuous-control tasks demonstrate improved robust control and unseen-task generalization, with and without disturbances.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.