acceptodds
Under review as a conference paper at ICLR 2027

Provably Variance-Reduced Federated Categorical Distributional TD Learning with Linear Function Approximation

Abstract

Federated distributional policy evaluation enables multiple clients to estimate the entire return distribution from locally generated interaction samples without pooling raw data, thereby combining uncertainty-aware evaluation with distributed computation. However, finite-sample analyses of linear categorical distributional TD have largely been restricted to centralized settings, while theoretical studies of federated TD primarily address expected returns. This leaves a gap in the understanding of variance reduction for distributed return-distribution evaluation with multiple local updates. We study federated distributional policy evaluation under linear categorical function approximation and propose FedVR-DistTD, which combines a global control direction aggregated from distributed control batches with same-sample TD differences. Under homogeneous clients, full participation, and independent generative-model sampling, We show that the final server iterate , obtained after communication rounds, converges geometrically in mean square to a statistical neighborhood determined by the finite control batch With a fixed per-round control batch, the tail-averaged server iterate converges to the projected fixed point in mean square at rate , where is the per-client local-iteration budget. This result yields a corresponding expected squared Cram\'er-error guarantee, and is supported by experiments on a four-state Markov reward process.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.