acceptodds
Under review as a conference paper at ICLR 2027

Distributed Natural Policy Gradients with Poissonized Determinantal Averaging

Abstract

Natural Policy Gradients (NPG) can exploit the low-rank structure of the empirical Fisher via the Woodbury identity, replacing a parameter-space solve with a sample-space system. While computationally attractive, this formulation incurs memory for samples and parameters. Data parallelism can reduce this cost, but naively averaging local empirical NPG directions does not recover the pooled-batch estimator due to finite-sample bias from Fisher inversion and Fisher–gradient coupling. We propose Poissonized Determinantal Averaging (PDA), a communication-efficient distributed NPG method based on Poissonized local sampling and determinant-weighted aggregation. PDA communicates only one parameter-sized direction and one scalar per worker. We show that, for any fixed number of workers , PDA matches the full leading-order bias term of pooled Poissonized NPG as the expected local batch size grows, leaving an gap in expectation. For fixed , PDA also converges to the population damped NPG direction as . Across eight MuJoCo and Procgen tasks, PDA more closely matches pooled NPG than uniform averaging while retaining favorable compute and memory scaling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.