KITE: SPECIALIZING FROZEN MULTIMODAL MODELS WITH A SHARED SIDE-INPUT EXPERT
Abstract
Specializing a multimodal model with external evidence raises a twofold design challenge: how to integrate that evidence without consuming the host’s input sequence, while preserving its behavior when specialization is inactive. We introduce Kite, a shared side-input expert that reads external evidence and host marker states at sparse depths and writes additive updates at those markers. Marker-free requests bypass the branch, preserving host outputs bit-exactly under the same execution configuration; a training-free, prompt-level router controls invocation. We instantiate Kite for Earth observation using frozen DINOv3 features and a frozen 35B host, training a 1.88B-parameter side expert. On a pre-locked new-site recognition test, Kite exceeds a host LoRA matched across the full training chain by +9.8 percentage points. On a separate pre-locked test, mismatched imagery reduces accuracy from 0.429 to 0.106, supporting evidence-sensitive recognition. Benchmark adaptation confined to the expert raises the official VRSBench-VQA ten-category mean from 0.627 to 0.772 and reaches 0.913 on RSVQA-LR. These gains incur additional side computation. We report its cost alongside controlled ablations and a pre-registered negative result on depth-structured evidence use.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.