Block-Causal Encoding: Converting Large MLLM Encoders for Efficient On-Device Inference
Abstract
Large pretrained vision encoders provide the detailed representations needed by Multimodal Large Language Models (MLLMs), but high-resolution encoding is costly on mobile Neural Processing Units (NPUs). Beyond its quadratic computation, their full attention couples image regions across layers, making it difficult to complete encoding in small blocks suited to limited on-chip memory. We introduce BLoCE (Block-Causal Encoding), which preserves the original ViT layer structure and standard operators while restructuring attention for blockwise execution. It features block-causal attention allowing each image block to complete all encoder layers without waiting for future blocks and a bounded historical key-value cache supplying context for subsequent blocks. This organizes encoding around smaller intermediate features and attention inputs while retaining every output token. We further design Progressive Conversion to adapt the pretrained encoder to the new attention pattern through post-training, avoiding encoder pretraining from scratch. Based on AIMv2-300M, BLoCE achieves a comparable performance across 12 MLLM benchmarks. On a MediaTek Dimensity 9500 NPU, it reduces encoder latency at from 2220 ms to 699 ms, a speedup. These results demonstrate the value of restructuring attention for efficient high-resolution inference while reusing large pretrained vision encoders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.