Encoded but Not Reported: Label-Free Consistency Audits of Longitudinal Vision-Language Models
Abstract
Vision-language models are asked how a disease changes between scans, but benchmarks score each question alone. We audit them without labels using a property of true change: change over an interval is the sum of changes over its parts. Answers obeying this cocycle rule form the gradient subspace of combinatorial Hodge theory, so a model's distance to it lower-bounds its error. Clipped at a pre-declared threshold, this bound gives a distribution-free cohort certificate. We also probe frozen representations and repair answers by projection onto this subspace. Experiments on brain MRI and satellite imagery show that most open-weight models fail to beat a no-change predictor. On a glioma cohort with large changes, representations across model families encode change that answers do not report, and label-free probes recover most of it. Fine-tuning on projected answers corrects their direction but not their magnitude. Code, results and model outputs are available at https://anonymous.4open.science/r/VLM-3j-ICLR27-725E/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.