Cells or Variant Families? Resampling-Unit Sensitivity in Repeated-Variant VLM Benchmarks
Abstract
Repeated-variant vision-language model (VLM) benchmarks score multiple prompts, option rotations, or counterfactual images derived from a shared source. Treating those cells as independent can change reported uncertainty while leaving the observed score, eligible cells, and weights unchanged. We conduct a version-pinned descriptive audit of 23 public response matrices from DynaMath, MMBench V11, and VB, comparing cell-IID and source-family percentile-bootstrap intervals. Median family-to-cell interval-width ratios are 2.329, 1.827, and 0.865, respectively; the MMBench result is a rotated-prompt cell diagnostic, not its official CircularEval score. Across ten fixed model pairs, one unadjusted pointwise comparison changes from positive to unresolved; the other nine are unchanged. A post-hoc Bonferroni analysis also changes one comparison, but a different one. We report a scorer correction to our earlier analysis and present a descriptive re-analysis rather than a prospectively verified confirmatory study. The contribution is a bounded VLM-specific cross-benchmark measurement, not a new bootstrap method or a field-wide prevalence claim. The results motivate declaring the target population, scoring and eligibility rules, and resampling unit, and testing sensitivity while keeping the observed statistic fixed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.