acceptodds
Under review as a conference paper at ICLR 2027

Cells or Variant Families? Resampling-Unit Sensitivity in Repeated-Variant VLM Benchmarks

Abstract

Repeated-variant vision-language model (VLM) benchmarks score multiple prompts, option rotations, or counterfactual images derived from a shared source. Treating those cells as independent can change reported uncertainty while leaving the observed score, eligible cells, and weights unchanged. We conduct a version-pinned descriptive audit of 23 public response matrices from DynaMath, MMBench V11, and VB, comparing cell-IID and source-family percentile-bootstrap intervals. Median family-to-cell interval-width ratios are 2.329, 1.827, and 0.865, respectively; the MMBench result is a rotated-prompt cell diagnostic, not its official CircularEval score. Across ten fixed model pairs, one unadjusted pointwise comparison changes from positive to unresolved; the other nine are unchanged. A post-hoc Bonferroni analysis also changes one comparison, but a different one. We report a scorer correction to our earlier analysis and present a descriptive re-analysis rather than a prospectively verified confirmatory study. The contribution is a bounded VLM-specific cross-benchmark measurement, not a new bootstrap method or a field-wide prevalence claim. The results motivate declaring the target population, scoring and eligibility rules, and resampling unit, and testing sensitivity while keeping the observed statistic fixed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.