acceptodds
Under review as a conference paper at ICLR 2027

Cola2: Co-Adaptive Hierarchical Language Modeling with Continuous Semantic Compression and Balanced Discrete Generation

Abstract

Hierarchical language generation separates continuous semantic modeling from discrete text realization, but its effectiveness depends on what information the latent space should compress and how the discrete generator should use it. We propose Cola2, which jointly studies continuous semantic compression, the discrete reconstruction–generation balance, and their co-adaptation. Cola2 dynamically segments text into semantic spans, controls the compression rate with an auxiliary-loss-free negative-feedback mechanism, and combines reconstruction and generation supervision to learn compact and predictive latent representations. A block-causal masked diffusion model (MDM) decoder is jointly trained under clean, noisy, and dropped latent conditions, and classifier-free guidance (CFG) adjusts the balance between semantic reconstruction and autonomous generation. During training, token-prediction gradients update the encoder through the latents, driving the continuous representation and discrete generation to co-adapt, while Flow Matching prior learning is decoupled from representation learning via stop-gradient. At inference, the generated text is re-encoded under the fixed span layout, keeping the latent history consistent with the actual text context. Systematic component analyses validate the core designs, and on five downstream tasks with about 210B training tokens, Cola2 exhibits competitive scaling potential, providing empirical support for hierarchical language generation with continuous–discrete co-adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.