acceptodds
Under review as a conference paper at ICLR 2027

CacheSep: Amortizing Text-Queried Sound Extraction Across Queries in Codec Latents

Abstract

Text-queried target sound extraction models, including recent ones that operate in the compact latent space of a neural audio codec, rerun a query-conditioned separator for every query, even though a recording is often queried repeatedly for different sources. Each request thus repeats work on a mixture that has not changed. We present CacheSep, a codec-domain extractor that runs a query-independent Transformer trunk over the mixture's DAC codec latent once and caches it, then answers each query with a shorter tail conditioned on a frozen CLAP text embedding. In the codec domain, decoding the latent back to audio dominates each query's cost, so we also train a lightweight decoder for the DAC latent, 19.37x faster than DAC's original decoder at comparable reconstruction quality, and share it across all codec-latent systems. We train and evaluate on 1,227 hours of speech, music, environmental, machine, and animal recordings from 25 public datasets, mapped to 89 sound classes. On 8,900 held-out mixtures, every cache split we evaluate stays within 0.11 dB SI-SDRi of the matched codec-domain baseline CodecSep at comparable training cost, while later queries run up to 2.10x faster end to end. Sparse experts in the tail add a small quality gain at a latency cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.