SEARCH WHAT YOU MUST, SOLVE WHAT YOU CAN: FEDERATED CONTINUOUS PROMPT ADAPTATION FOR BLACK-BOX VLMS
Abstract
Federated adaptation of pretrained vision–language models (VLMs) commonly relies on lightweight prompts, yet existing approaches typically require gradient or feature access to the frozen backbone. Such white-box access may be unavailable when VLMs are exposed through restricted inference interfaces or opaque runtimes, while distributed users still need to adapt them to private local data. We therefore study federated adaptation of black-box VLMs, where clients can interact with the model through continuous multimodal prompts and class-score feedback but cannot access model internals. Rather than treating black-box adaptation as a single optimization problem, we ask a simpler question: what must be searched, and what can be solved? We propose FedMUSE, which decomposes federated black-box VLM adaptation into two complementary components: multimodal prompt search for what must be optimized through queries, and global decision refinement for what can be solved explicitly from the returned score space. FedMUSE first introduces mean–contrast prompt search, which parameterizes textual and visual prompts through common and differential latent kernels and performs block-wise zeroth-order estimation to reduce cross-block interference. It then freezes the selected prompt predictor and performs global residual refinement in the resulting shared score space. Clients communicate additive sufficient statistics rather than locally fitted heads, allowing the server to recover the pooled residual solution exactly without collecting individual scores or labels. Across 13 classification benchmarks, FedMUSE achieves 70.15% mean accuracy, improving its prompt-only predictor by 5.16 percentage points while adding less than 0.34% training queries, with compact communication and single-call VLM inference. Code is available at https://anonymous.4open.science/r/FedMUSE-04C1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.