Models Are Mortal, Measurements Are Not: A Portable, Model-Agnostic Representation for Video Misinformation
Abstract
A representation belongs to the model that produced it: replace the model and everything fitted on it must be rebuilt, and a closed model exposes no representation at all. We ask whether the coordinates of a representation can instead be defined in language, as a fixed bank of yes/no questions that any model can answer, and test this on short-video misinformation across eleven frozen vision-language models, two prompting languages, a closed API and human annotators. A linear readout fitted once on any one open model deploys unchanged on the other ten, losing only .040 AUC on average, and, with same-event corroboration, comes within a point of a fine-tuned detector's accuracy. A matched control separates two properties, compatibility and validity. A bank of questions unrelated to the task also transfers across models, yet loses much of its discrimination across datasets, so fixing the coordinates buys compatibility while their content buys validity. A measurement view splits each reading into what the question measures and how reliably a model reads it; it accounts for why aggregates transfer when single questions disagree, and locates where reading thins, above all in the smallest models and in collapsed answer channels. In simpler terms, the questions hold the representation, and a model is only one way of reading it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.