When Aggregate Accuracy Does Not Identify Functional Change: Sharp Bounds on AV-LLM Adaptation Effects
Abstract
Fine-tuning Audio-Visual Large Language Models (AV-LLMs) is commonly evaluated by in-domain and cross-domain joint accuracy, measured with both audio and video available. Yet the same reported accuracies can conceal different functional changes. Even audio-only, video-only, and joint accuracies omit whether the two isolated inputs succeed on the same questions. We define isolated-path coverage as the fraction of questions answered correctly by at least one single-modality input. Conversion is joint accuracy minus that coverage. From before-and-after aggregate accuracies, we derive a sharp identified set for the conversion change and establish exactly when its sign is identifiable. The joint-minus-video gap combines conversion with the fraction of questions answered correctly with audio alone but not video alone. Its change can therefore have the opposite sign to the conversion change. Our audit spans 56 checkpoint–dataset evaluations and 40 before-and-after fine-tuning comparisons across Qwen2.5-Omni, Qwen3-Omni, MiniCPM-o, and VITA. Aggregate reports identify none of the 40 conversion-change signs, while the conventional gap points in the opposite direction in 17 comparisons. Paired per-question evaluations recover the missing overlap. They reveal a cross-domain conversion reversal under the same adaptation and a coverage gain offset by a conversion loss behind a nearly unchanged joint score. The balance between coverage and conversion varies across adaptation datasets and model families. AV-LLM adaptation calls for paired functional profiles alongside joint accuracy to establish what changed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.