acceptodds
Under review as a conference paper at ICLR 2027

CARVE: Cascaded Adaptive Routing of Visual-Condition Embeddings for Multi-Reference Generation

Abstract

Multi-reference controllable generation represents each reference with many visual-condition tokens. As the number of references grows, these tokens create substantial attention cost in existing transformer architectures. Yet many tokens contribute little to the current generation while still being processed at every layer and denoising step. Existing compression methods use fixed, hand-crafted rules to select or merge tokens. Such rules do not adapt to the current input or generation state and can remove reference-specific identity, texture, or color details. We propose CARVE, the first learnable and input-adaptive framework for visual-condition-token compression. Rather than predefining token importance or pruning locations, CARVE learns to preserve the tokens that contribute to the current generation at every denoising step and to remove tokens after their reference information has become redundant. The resulting compression adapts to each input and denoising computation while preserving the information required for faithful multi-reference generation. Across diffusion-transformer and unified multimodal generation backbones, CARVE achieves the strongest generation quality among compared compression methods while substantially reducing token processing and inference latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.