acceptodds
Under review as a conference paper at ICLR 2027

Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study

Abstract

Text-output scores alone do not show whether quantization preserves performance when the transcript does not reveal the target label. We evaluate fixed 6- and 7-bit Qwen2-Audio-7B-Instruct allocations on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The provenance-bound FLEURS rerun has selected-minus-FP16 BLEU and chrF intervals that all include zero. On RAVDESS, where the same two sentences occur equally often with every emotion label, canonical-label accuracy is -3.71% for 6 bit and -1.17% for 7 bit relative to FP16. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows that a translation score does not establish performance on a speech task whose target cannot be recovered from the transcript.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.