Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models
Abstract
Large audio language models (LALMs) increasingly reason over long audio clips, motivating audio compression to reduce memory use and inference latency. However, audio compression can leave the model's overall answer accuracy acceptable while severely degrading accuracy for particular query families. We introduce a framework for auditing a given audio compression method by measuring the increase in an LALM's answer error relative to uncompressed audio. We formalize an acceptance criterion that limits this increase for every required query family and derive a practical protocol for testing compression settings with statistical confidence. Experiments spanning five multiple-choice audio question-answering benchmarks and four LALMs expose compression-induced damage hidden by overall accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.