SAGE: Efficient Protected Vision-Language Inference on Smart Glasses
Abstract
Smart glasses are a promising platform for vision-language assistants, but private split inference remains expensive: protecting visual representations before transmission does not remove the cost of on-device vision encoding, wide protected tokens, or hardware-unaware execution. We present SAGE Smart-glasses-Aware Guarded Encoding), a hardware-aware framework for efficient protected VLM inference on smart glasses. SAGE keeps raw images and unprotected visual tokens on the trusted device, reduces visual input cost, executes shape-specialized INT8 graphs, and compresses protected representations into compact 256-dimensional INT8 packets before transmission. Unlike efficiency methods optimized only for nominal model size, SAGE selects its operating point using downstream answer quality together with measured latency, memory, and thermal behavior on real Mentra Live glasses. On a matched Qwen workload, SAGE reduces glasses inference latency by 39.18%, device-sensor temperature rise by 21.91%, and protected packet size by 96.56%, while retaining 85.62% of the protected reference's relaxed utility and comparable prompted containment. Direct glasses-to-phone tests deliver all 192 packets and reduce median packet-plus-acknowledgment time from 71.3 to 10.0 ms. The compact protected interface also transfers across four additional VLM families. These results show that private wearable VLM inference must be optimized jointly across representation protection, communication, and real-device execution rather than through model compression alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.