acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Share: Latent Representations for FedRL with Observation Distractors

Abstract

Federated reinforcement learning lets agents in different environments learn a shared policy by averaging updates without exchanging trajectories. This setting becomes harder when each observation contains task-irrelevant signals: the policy must learn from returns which coordinates matter, and each client learns that lesson from a different noise realisation. We study whether the federation should share a learned representation rather than raw-observation policies. We formalise the setting as a federated exogenous block MDP, then train federated PPO with an encoder whose gradient is routed through the value loss, with or without a self￾supervised objective. The objectives include a JEPA-style multi-step predictor, a behavioural metric, and a multi-step inverse model. A linear–Gaussian analysis shows why a noise-free representation is common across clients and why policy￾gradient learning on raw observations pays a fixed-step-size penalty that grows with the number of distractors. On six MuJoCo tasks with clients that differ in physical parameters, the critic-trained encoder outperforms FedAvg, FedProx, and FedKL baselines under 64 appended Gaussian distractors. The gain is largely due to gradient routing rather than model width. Self-supervised objectives help on some tasks and hurt on others. For the predictive agent, the gain grows with the number of i.i.d. distractors, but its size changes with temporal correlation: it is reduced on several tasks at ρ = 0.99 and remains positive on most. Federation improves the critic-trained encoder on all four tasks with a single-agent comparison, although it is not the source of the basic advantage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.