FEDWM: PERSONALIZED FEDERATED LEARNING FOR HETEROGENEOUS VISUAL WORLD MODELS
Abstract
Federated learning can combine experience across clients to train visual world models without pooling their trajectories. When clients differ in appearance and dynamics, the central design question is which parts of the predictor to share and where to personalize. We introduce FedWM, a federated training framework that places private residual modules at both the interfaces and internal layers of action-conditioned world models. Clients first learn a shared core through parameter averaging, then freeze it and train private modules locally with the backbone's native prediction objective. We instantiate FedWM separately with predictive DINO-WM and diffusion-based NanoWM across PushT, PointMaze, Wall, and Rope. With matched training steps, FedWM lowers mean held-out prediction error relative to interface-only personalization in all eight backbone–environment settings, each evaluated with three training seeds. A near-capacity-matched DINO-WM PushT ablation also favors internal residuals over wider interfaces. In DINO-WM studies, FedWM retains its advantage over interface-only adaptation with eight clients at fixed total data and in most low-data new-client settings. Closed-loop PushT comparisons favor FedWM over FedAvg on validation but not on the frozen test bank, highlighting the need to evaluate prediction and control separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.