Adaptive Layerwise Perturbation for Stable Off-Policy LLM Reinforcement Learning
Abstract
Training large language models with reinforcement learning often reuses rollout batches and generates them with a separate inference engine. These practices create policy staleness and training–inference mismatch, which can destabilize importance-weighted updates. We introduce Adaptive Layerwise Perturbation (ALP), which learns a Gaussian perturbation scale for each layer and samples hidden-state perturbations during policy updates. The perturbed training policy forms the numerator of an importance ratio against the unperturbed rollout policy; rollout generation and evaluation remain unchanged. A smoothed-policy analysis relates perturbation to mismatch sensitivity and objective curvature. Across single-turn math and multi-turn tool-integrated reasoning with three model backbones, ALP achieves the highest average accuracy among the evaluated methods. In the five-seed multi-turn comparison, Seq-ALP reaches average accuracy versus for the strongest baseline, Token-MIS. ALP also shows controlled KL divergence, smaller importance-ratio tail excursions than Bypass, and higher multi-turn pass@ at fixed rollout budgets. Location ablations favor perturbing all layers over the tested partial-layer and logits-only variants.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.