acceptodds
Under review as a conference paper at ICLR 2027

Fast Adversarial Implicit Diffusion for Model-Based Reinforcement Learning

Abstract

Adversarially guided trajectory diffusion aims to improve robustness in model-based reinforcement learning by steering a diffusion world model toward low-return trajectories, but this guidance is computationally expensive. We find that diffusion sampling accounts for 80.3% of training wall-clock time. Accelerating it is nontrivial: shortening the denoising chain changes the allocation of the method's risk budget, while deterministic sampling causes the transition covariance required by adversarial guidance to degenerate. We introduce AID-RRL (Adversarial Implicit Diffusion for Robust Reinforcement Learning), a training-free sampling modification that uses a strictly stochastic DDIM-style implicit sampler over a strided sequence of denoising steps and recalibrates the risk budget to the steps actually taken. Across nine independently trained MuJoCo Hopper checkpoints, AID-RRL preserves rollout fidelity within a pre-registered equivalence margin while reducing per-batch trajectory generation cost by 6.0–12.3×. In single-seed end-to-end runs, it reduces training time by 2.7–3.1×. A second pre-registered test reveals a more fundamental limitation: neither AID-RRL nor the original full-chain method measurably shifts generated returns toward the intended risk percentile. We show that this limitation is structural rather than a consequence of acceleration: the adversarial displacement scales as O(σ²) while diffusion noise scales as O(σ), causing guidance to weaken relative to noise as σ decreases. Thus, AID-RRL removes a major computational bottleneck while exposing a general diagnostic for risk-aware diffusion guidance: its displacement scale should be compared with the noise it must overcome before expensive training is undertaken.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.