Attacking and Defending One-Step Flow Policies in Offline Reinforcement Learning
Abstract
One-step flow policies enable efficient action generation in offline reinforcement learning, but remain vulnerable to observation perturbations. We introduce Behavior-Decoupled Adversarial Perturbation (BDAP), which trains a generator against a frozen surrogate to lower estimated action value at the true state while increasing action deviation from a behavior-only reference. At deployment, BDAP generates perturbations in one forward pass without victim queries or gradients. The existing Diffusion Model-Based Predictor (DMBP) partially mitigates these attacks but leaves reconstruction errors. We therefore propose the Residual-Corrected Filter (RCF), which predicts these residuals from the perturbed observation and the DMBP-reconstructed observation and corrects the reconstruction before policy execution. RCF is trained only on simulator rollouts with uniform random observation perturbations, keeping the policy and DMBP frozen, and requires no attack-type information at deployment. Our analysis bounds reconstruction error through prediction loss, target clipping, and correction constraints, and connects the remaining error to a return-gap bound relative to the clean policy under Lipschitz assumptions. Experiments on OGBench and D4RL across four policy families demonstrate effective cross-policy transfer of BDAP and aggregate performance gains from RCF over DMBP under adversarial attacks and Gaussian noise.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.