From Visual Token Reduction to Realized Edge Efficiency: A Stage-Wise Study with CIVIC
Abstract
Reducing visual tokens is a common strategy for accelerating vision–language models, yet token counts and operation counts do not establish faster batch-one inference on edge hardware. Compression after visual encoding leaves expensive backbone layers untouched, while early compression may be undone at the multimodal interface or offset by aggregation overhead. This paper presents Compact Inference for Vision-Language Integrated Compression (CIVIC), which aggregates visual states at an intermediate encoder layer and carries the compact representation through the remaining visual blocks, projector, and language-model prefill. A configurable visual budget controls aggregation, and text-aligned distillation transfers predictions across unequal visual-token sequences. The evaluation adapts CIVIC to Qwen3-VL-4B-Instruct and profiles batch-one inference on Jetson AGX Orin, tracing interface lengths and timing aggregation, the encoder, the projector, and prefill separately. The median final prefix falls from 816 to 470 tokens; derived response-length-standardized latency decreases by 26.6%, with lower measured encoder and prefill stage times. In a fixed-configuration control, restoring dense downstream tokens increases derived latency, whereas optional visual key–value pooling adds device cost. CIVIC reduces derived latency on the four scored benchmarks and within each Video-MME duration group. Task scores show benchmark-dependent quality trade-offs. Together, the stage traces and fixed-budget controls connect persistent compactness with realized edge acceleration despite the measured aggregation overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.