acceptodds
Under review as a conference paper at ICLR 2027

2D Learned Speech Compression with Content-Adaptive Time-Frequency Modeling

Abstract

Neural audio codecs have relied on one-dimensional sequence representations, where speech is transformed into temporally ordered latent representations. However, this formulation does not explicitly model the inherent time-frequency structure of speech, treating frequency as a representation-dependent dimension and limiting natural adaptation across different sampling rates. To address this core limitation, we represent speech as a time-frequency spatial field via a complex short-time Fourier transform, preserving frequency and time as explicit spatial dimensions. Inspired by learned image compression, we then formulate speech coding as a two-dimensional learned compression problem over this spatial field, applying learned image compression models directly to compress the field. This enables a single model to operate on native 16, 24, and 48 kHz speech without sampling-rate-specific branches. We further apply content-adaptive token organization to the time-frequency field for sequence modeling, stabilize entropy skipping, and merge entropy-coded streams into a compact bitstream. Experiments show that CA2D, with only 69M parameters, achieves strong low-bitrate quality while maintaining a unified architecture across sampling rates, a notable achievement given its compact size. This single compact model operates across 16, 24, and 48 kHz and consistently provides competitive or superior objective and subjective quality compared with existing neural codecs: at around 800 bps, CA2D achieves comparable or better quality than 1 kbps-class models with over 100M parameters, such as BigCodec (159M parameters, 1040 bps), while the same model at 48 kHz is comparable to the 6 kbps operating point of FlowDec, a full-band codec. These results clearly demonstrate the effectiveness of treating speech as a two-dimensional spatial field for scalable neural audio compression with native multi-rate operation and low-bitrate perceptual quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.