What Survives Extreme Compression? Vision-Language Embedding Alignment at Microcontroller Scale
Abstract
We ask what remains of CLIP’s image - text embedding geometry when ViT-B/32’s 87.8 M-parameter image tower is distilled into a 724 K-parameter microcontroller encoder whose class prototypes the teacher’s text encoder computes offline. Distilling on CC3M, we find a sharp capacity - granularity boundary: the student keeps 39% of teacher accuracy on COCO with classes weighted equally, 12% on ImageNetV2, 5 - 9% on fine-grained sets and near chance on two robotics sets, a boundary that tracks pretraining coverage. It spans ≈93 effective dimensions where the teacher needs ≈277, and a stronger teacher adds up to 3.2 points at the deployed dimension at no device cost. One coordinate of the teacher is its prototypes’ shared direction: CLIP ViT-B/32’s coordinate 133 carries 79% of the mean prototype’s squared norm and enters the prefix only at 256 dimensions, where it adds ≈6 per-image COCO points, all on “person”, while class-balanced accuracy falls; projecting it out lifts class-balanced 256d above 128d on COCO and ImageNetV2. Alignment is also teacher -specific: prototypes from any other text encoder score at chance, but a closed-form map fitted on class names alone recovers 64-71% of accuracy offline. The INT8 student runs at 2.5 FPS on an ESP32-S3 and fits an STM32H7 in 1,085 KB flash (464 KB arena, emulated); its 128 px input costs eleven points of top-1 at ten COCO classes (37.9% vs. 48.8%).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.