acceptodds
Under review as a conference paper at ICLR 2027

Single-Timescale Personalized Average-Reward Federated TD Learning Across Heterogeneous Environments

Abstract

We study personalized federated TD learning, in which a collection of agents interact with different environments and jointly learn their respective value functions. Inspired by the recent success of personalized federated learning (PFL), we focus on the setting where there exists a shared linear representation and the agents' optimal weights collectively lie in an unknown linear subspace. Under canonical single-timescale updates with average-reward objectives and Markovian sampling, we propose a natural personalized federated TD algorithm, named \bf pFedAR-TD, in which agents jointly estimate a common subspace while updating agent-specific weights and mean rewards. We show that this decomposition can filter out conflicting signals, effectively mitigating the negative impacts of “misaligned” signals arising from environmental heterogeneity, and hence achieve linear speedup. Compared with the discounted-reward setting, the combination of personalization and average-reward objectives significantly complicates the convergence analysis. In particular, the additional need to learn local mean rewards results in three tightly coupled dynamics rather than two, calling for a careful balance between fine-grained characterization of error perturbations and the overall tractability of the system-level analysis. Our analysis exploits geometric properties induced by key algorithmic components such as projection, innovation filtration, and QR decomposition. A large subspace error early in training can amplify perturbations in the coupled iterates and jeopardize convergence. To address this, we further introduce an initialization phase and establish a maximal-deviation bound to keep the coupled iterates in a region where the convergence argument applies. Experiments are provided to show the benefits of learning via a shared structure to the more general control problem.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.