VLM DNA: Tracing Vision-Language Models via Distance-Preserving Representations
Abstract
Identifying vision-language models (VLMs) and tracing their relationships across fine-tuning trajectories requires compact model representations that can be ex- tracted without access to model internals. We study training-free VLM identity representation extraction using only image-question queries and generated re- sponses, without modifying model parameters or requiring shared architectures. Existing text-only response representations do not explicitly account for the rela- tionship between generated answers and their visual inputs, potentially overlook- ing useful signals for distinguishing VLMs. We introduce VLM DNA, a compact representation of model behavior, and DNAExtract, a training-free framework for extracting it. DNAExtract probes each model with a shared set of image-question pairs and encodes the images and generated responses using frozen CLIP image and text encoders. It then constructs outer-product interaction matrices to repre- sent image-response associations and applies sparse random projection to obtain compact DNA vectors designed to approximately preserve pairwise distances be- tween the original cross-modal representations. Experiments on 39 open-source VLMs from seven families demonstrate that VLM DNA achieves an AUC of 0.929 for relation detection, an AUC of 0.923 for one-to-one verification, and a rank-1 accuracy of 89.38% with a mean average precision of 0.936 for one-to-many identification. These results outperform text-only response embedding baselines and support the utility of cross-modal representations for VLM identification and trac- ing. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.