acceptodds
Under review as a conference paper at ICLR 2027

FIRA: DEPLOYMENT-TIME RESCUE OF FROZEN RL AGENTS VIA A LOGICAL FIREWALL

Abstract

Frozen reinforcement-learning policies can fail at deployment in narrow, recognizable ways, yet retraining them may be impractical. We study deployment-time rescue: correcting a frozen policy with a separately trained intervention layer, without modifying the policy or adapting during deployment. Training such a layer offline exposes an offline-to-live rescue gap: models with similar offline accuracy can behave very differently in closed loop, and running the rule-based supervision procedure as a live controller does not predict how well the learned models will perform. We introduce FIRA (Firewalled Intervention via Relational Atoms), which represents scenes as ego-centric graphs of affordance-tagged entities and predicts atoms describing threats, resources, and progress. A logical firewall blocks action-loss gradients through the predicted-atom pathway and permits the rescue action to override the frozen policy only when an intervention atom fires. FIRA is trained on fixed offline scenes labeled from environment rules, without reward optimization, rescue-policy rollouts, or deployment-time adaptation. Under identical supervision, matched flat-MLP and behavior-cloning baselines attain high offline accuracy but transfer less consistently to closed-loop control. Across three Atari games, two frozen-policy families, and two ViZDoom scenarios, FIRA rescues novel hazards, resource mismanagement, reactive-control failures, unseen threats, and health depletion. On Frostbite, FIRA raises the frozen policy's score from 126 to 2527 and advances beyond the first level in 19 of 20 episodes, compared with 0 of 20 without rescue.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.