acceptodds
Under review as a conference paper at ICLR 2027

pFedVSA: Personalized Federated Visual-to-Semantic Alignment for CLIP Adaptation

Abstract

Federated few-shot CLIP learning aims to exploit client-specific visual evidence for personalization while preserving transferable recognition over global base classes and unseen novel classes. A key challenge is how to use client-specific class-level visual evidence for collaborative adaptation without directly sharing such evidence, especially when local visual priors can be sparse or unreliable. We propose Personalized Federated Visual-to-Semantic Alignment (pFedVSA), which separates private visual personalization from shared semantic adaptation. Each client maintains a personalized image adapter and a Local Visual Prior Memory (LVPM) that remain local throughout federated training. A shared Prototype-to-Semantic Transfer (PST) module learns to map locally available visual priors into CLIP-compatible semantic pseudo-tokens, enabling collaborative visual-to-semantic adaptation without aggregating the priors themselves. Semantic Alignment in CLIP Space (SACS) further incorporates prototype-conditioned evidence through reliability-controlled residual calibration based on prior support and image–prior consistency. For classes without local visual support, prediction falls back to the standard adapted semantic branch. Extensive experiments on eleven datasets under label and feature heterogeneity show that pFedVSA consistently improves the balance among client-specific recognition, global base-class transfer, and novel-class generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.