acceptodds
Under review as a conference paper at ICLR 2027

TimePrune: Adaptive Temporal Resolution for Efficient Audio Large Language Models

Abstract

Audio representations form temporally ordered token trajectories whose variation can differ substantially over time. Processing dense audio tokens throughout Audio Large Language Models (ALLMs), however, incurs substantial computational overhead. Existing audio token compression methods address this inefficiency by exploiting token importance, redundancy, or temporal structure, but typically do not directly assess how faithfully the retained tokens preserve the underlying representation trajectory. We introduce TimePrune, a training-free framework that preserves temporal representation trajectories by adaptively determining temporal resolution under a given token budget. At its core, Time-Aware RDP recursively refines intervals whose intermediate representations are insufficiently described by their temporal anchors, yielding adaptive local anchor density according to trajectory deviation. We further introduce Temporal Coverage Refinement to regulate excessively sparse temporal intervals while preserving trajectory-driven refinement. Extensive experiments across four ALLMs and five audio understanding benchmarks demonstrate the effectiveness of TimePrune across different models and token budgets. In our efficiency analysis, TimePrune retains only 15% of audio tokens while preserving 98.4% of the Vanilla score and reducing audio-token-dependent LLM computation by 85.2%, demonstrating its effectiveness for efficient ALLM inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.