Aggregation-Faithful Value Adaptation for Federated Vision-Language Models
Abstract
Federated adaptation of large vision-language models (VLMs) is appealing for privacy-sensitive applications but must contend with client data heterogeneity, strict communication budgets, and on-device inference constraints. Recent methods predominantly adopt lightweight prompt or adapter interfaces, yet averaging their parameters may fail to preserve the functional changes learned across heterogeneous clients. We therefore ask: can robust federated VLM adaptation be achieved through an aggregation-faithful intrinsic interface? We answer this question with TriVA (Tri-State Value Adaptation), a module-free framework that fine-tunes only native value biases in both visual and textual encoders. TriVA operates on existing backbone parameters and introduces no additional inference-time modules. Our analysis shows that this intrinsic interface provides substantially higher aggregation fidelity than the tested prompt- and adapter-based alternatives. To make this aggregation-friendly interface effective under heterogeneous clients without excessive synchronization, we further introduce a one-shot calibration mechanism that assigns value-bias groups to three persistent states: Frozen groups preserve pretrained values, Shared groups participate in federated aggregation, and Private groups retain client-specific adaptation locally. The resulting assignments remain fixed throughout training, with only shared groups exchanged for aggregation. Extensive experiments across heterogeneous federated benchmarks demonstrate strong adaptation performance with a compact trainable parameter set and substantially reduced communication cost. Semantic segmentation experiments further demonstrate the applicability of TriVA beyond image classification to dense prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.