acceptodds
Under review as a conference paper at ICLR 2027

GAP: Generation-Aligned Paired Fidelity Evaluation for GGUF-Quantized VLMs

Abstract

Reliably distinguishing small fidelity differences between quantized candidates is challenging during quantization development. We introduce GAP, a Generation-Aligned Paired fidelity evaluation protocol for GGUF-quantized vision-language models. GAP generates image-conditioned responses with the full-precision reference, teacher-forces all candidates on these shared responses, and applies item-level paired tests to their fidelity divergence (e.g., KLD) from the reference to detect small fidelity differences with quantified uncertainty. A pilot-based procedure selects mode-specific token caps to reduce per-item evaluation cost, enabling larger samples and greater statistical power within a fixed budget. Across 126 quantized checkpoints near the Q4 regime, six benchmarks, and 15 model-mode configurations, the median paired-to-unpaired SNR ratio is 3.7 in thinking mode and 2.0 in non-thinking mode for candidate pairs with relative fidelity gaps below 5%; pairing reduces the estimated median sample size per candidate required for 80% power by approximately 15.5× relative to an independent-item unpaired design for such pairs in thinking mode, whereas close non-thinking pairs still need an estimated 11,000 items, beyond our budget. Fidelity rankings can also differ between thinking and non-thinking modes, while responses generated by a different reference model can change verdicts, including statistically significant reversals. These findings support evaluating candidates in their intended inference mode using responses from their corresponding reference model. We release the associated artifacts on Hugging Face Hub: https://huggingface.co/datasets/Anonymous778862/anony778862cr.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.