Scalable Feature Coding for Distributed Large Vision-Language Model Inference
Abstract
Feature coding enables distributed inference of large vision-language models (LVLMs) by compressing visual representations at the edge and transmitting them to the cloud for downstream processing. However, using a single coded representation for tasks with different information requirements can waste bandwidth and incur unnecessary computation in the cloud-side large language model (LLM). To address these problems, we propose the first Scalable Feature Coding (SFC) framework for LVLMs. SFC organizes visual representations into a hierarchical bitstream, comprising a base layer that encodes coarse-grained global semantics and two enhancement layers that provide additional information for object-level and pixel-level semantics, respectively. This hierarchy enables coarse-grained tasks to operate on fewer visual tokens while supplying finer-grained information on demand. To further improve coding efficiency, we analyze correlations across and within information layers and introduce two complementary context models. Cross-layer context modeling (CL-CTX) conditions the entropy model of each enhancement layer on previously decoded layers, exploiting inter-layer dependencies for more accurate probability estimation. Intra-layer context modeling (IL-CTX) further reduces coding redundancy by capturing dependencies within a layer. Experiments with InternVL on image captioning, object detection, and instance segmentation demonstrate that SFC achieves 18.37% – 51.30% BD-BR and 7.5% – 55.6% end-to-end LVLM inference latency over baseline methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.