Federated Learning of Large Novel View Synthesis Model for Better Personalized Neural Field
Abstract
Large image-conditioned novel view synthesis (NVS) models benefit from training across diverse scenes, but collecting multi-view images from many owners on a central server can be impractical due to privacy and data-ownership constraints. Keeping each scene local avoids raw-data sharing, yet personalized neural fields trained from sparse client views often suffer from insufficient geometric supervision. Moreover, sharing such neural fields is itself undesirable because their weights can constitute a directly renderable representation of the private scene. We propose a federated framework that couples a server-side large NVS model with personalized client neural fields that remain entirely local. Each client uses its neural field to generate nearby pseudo-target views and locally refine only a trainable subset of the server model using both captured and rendered supervision. Only the resulting parameter update is communicated, while raw images, pseudo-targets, and neural-field parameters remain local. The aggregated global model is then returned to clients to improve their neural fields through consistency supervision, creating a mutual-reinforcement loop between global generalization and local personalization. Experiments with 5,000 scene-level clients show substantial improvements in personalized client neural fields, while the server model achieves competitive generalized NVS performance under zero-shot, cross-domain, and non-IID evaluations. Empirical MIA/PIA and limited-context leakage evaluations further indicate lower privacy exposure than directly sharing client neural fields, without claiming formal differential privacy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.