COMPACT: A UNIFIED FRAMEWORK FOR MULTIMODAL SPLIT LEARNING WITH COMPRESSED ACTIVATIONS
Abstract
Split learning enables collaborative training of vision–language models without sharing raw data, but transmitting intermediate activations across distributed clients can incur substantial communication overhead. We investigate a lightweight multi-client setting in which client-side modules are frozen, eliminating gradient transmission and parameter synchronization, and focus on activation compression at the client-server boundary. We introduce COMPACT, a unified framework that supports heterogeneous vision–language backbones with interchangeable compression modules, and evaluate seven compression strategies across four vision–language models on four medical VQA benchmarks. We find that compression effectiveness is strongly model-dependent, and blockwise NormalFloat activation quantization (Ql) provides the most consistent performance across backbones. Amongst compression configurations, larger cosine distance from uncompressed representations is consistently related to downstream accuracy, with Spearman correlations ranging from −0.72 to −0.85. Layer-wise analysis further shows that internal representation drift caused by reconstruction errors accumulates through frozen layers. Moreover, we identify two recurring phenomena associated with performance degradation: unstable training of reconstruction modules under short training budgets and poor utilization of quantization levels. Finally, in a simulated cross-silo setting, COMPACT reduces communication by 72.5% for 4-bit Ql, while preserving 98.5% baseline performance. Together, these results show that compression effectiveness depends on both compression rate and its interaction with pretrained representations and downstream adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.