acceptodds
Under review as a conference paper at ICLR 2027

LayerBridge: Unifying Understanding, Generation and Editing with Adaptive Attention over Frozen VLM Layers

Abstract

Recent approaches to unified multimodal modeling use pretrained vision-language models (VLMs) to encode prompts or noisy visual representations as conditions for a generation decoder. However, these methods typically condition the decoder on final-layer features or on fixed combinations of multi-layer prompt features, limiting each decoder block's ability to select the layers it needs as denoising progresses. Motivated by this, we propose LayerBridge, which uses depth attention to let a pixel-diffusion decoder dynamically select and combine features from multiple layers of a frozen VLM's vision encoder and language model, reusing its learned representations while obtaining multi-granularity guidance. The decoder uses its current state and the denoising timestep to weight and combine these features, allowing different mixtures across blocks, image regions and noise levels. This adaptive readout also provides a general architecture for feature reuse across generation, understanding, and editing. For understanding, the same mechanism lets visual tokens within the VLM's language model query features from multiple vision-encoder layers, providing finer visual information for vision-centric perception. For editing, reference tokens retrieve fine visual cues from intermediate vision layers and pass them to target states through LLM attention, preserving reference details during editing. To learn this shared interface, we keep the VLM frozen while training the added modules across generation, understanding, and editing. With Qwen3.5-0.8B/4B and 0.3B/1.1B generation-specific parameters, LayerBridge reaches 0.87/0.90 on GenEval, 0.54/0.62 on WISE, and 3.86/4.16 on ImgEdit, together with gains in vision-centric perception over native Qwen3.5.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.