Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?
Abstract
Small vision-language models (VLMs; B parameters here) are attractive for edge deployment but difficult to adapt with scarce supervision and fixed weights. External tools and memory can help, but using them introduces decisions that the model may not execute reliably. We introduce Harness Compilation (HC), an offline procedure that uses a large teacher to build a reusable harness for a frozen student. The teacher revises task knowledge and control logic from student execution traces, and a held-out validation set selects the harness. Deployment requires neither weight updates nor teacher calls. We organize runtime control into five operational degrees of freedom: invocation, selection, argument generation, evidence integration, and stopping or abstention. Interventions show that activation, open argument generation and evidence selection can be costly for the student, while several content rewrites and output constraints have little effect. Across seven visual question-answering settings, HC improves task scores by 12.0 to 22.4 points over bare models. Useful fact cards transfer to all ten evaluated students, but control policies are model-dependent and may need recompilation. With 100 practice items, HC achieves higher mean scores than answer-only LoRA on InfoSeek, DocVQA and SlideVQA. At larger budgets, weight adaptation can match or exceed the harness, while combining the two improves SlideVQA beyond either route alone. These results support adapting the division of work to the student rather than uniformly removing its decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.