ScoreQA: A Score-Grounded Benchmark for Reading and Reasoning over Symbolic Scores
Abstract
Evaluating score-reading ability requires questions whose answers are determined by the supplied notation and explicit musical conventions. ScoreQA assesses defined Reading and Reasoning abilities on short ABC 2.1 score windows. Reading identifies events, notation attributes, and timing; Reasoning compares or transforms this information. A MusicXML-based oracle verifies answers and their evidence in ABC. Each quartet fixes one question and four choices across four valid windows supporting different answers. ARJQ describes item, paired-score, and complete-quartet correctness, together with obsolete-answer retention. Seven open-weight backbones evaluated with plain-text five-shot restricted next-token scoring achieve no complete Reasoning quartets. In a separate five-shot API evaluation with generated answer letters, five models achieve 92.8% to 99.9% Accuracy and 76.8% to 99.6% Quartet, including 72.0% to 100.0% Reasoning Quartet. Results are reported separately for the two prompting and inference settings. Supervised adaptation provides a supplementary use of the training split, raising Qwen3.5-9B’s Reading and Reasoning Quartet to 84.0% and 44.0%. ScoreQA provides a verifiable test of specified score-reading and analysis abilities, for which complete correctness is an intended outcome.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.