FactMix: Evaluating and Improving Mixed-Fact Robustness in Large Audio-Language Models
Abstract
Large Audio-Language Models (LALMs) achieve strong performance across speech, environmental sound, music, and audio reasoning tasks, yet they can overlook a single incorrect factual claim when it appears together with several correct claims in the same query. To study this failure, we introduce FactMix-Bench, a diagnostic benchmark built from paired Grounded Queries and Mixed-Fact Queries. Each pair shares the same audio and supporting facts, while the Mixed-Fact Query changes exactly one target fact to a plausible but unsupported alternative. FactMix-Bench covers source, content, attribute, temporal, and relation information, with 55,952 multiple-choice questions forming 27,976 pairs over 5,000 recordings. We further construct FactMix-Align, a preference dataset using the same controlled factual changes, with directionally balanced ordered perturbations to reduce training bias. Direct Preference Optimization with FactMix-Align improves Qwen3-Omni pair accuracy from 54.33 to 76.83, achieving the best overall performance among the evaluated models, while performance on established public audio benchmarks is largely preserved or modestly improved. These results show that strong general audio understanding does not necessarily imply reliable verification of multiple factual claims. The data and code associated with FactMix-Bench and FactMix-Align will be publicly released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.