acceptodds
Under review as a conference paper at ICLR 2027

Understanding Unified VLMs: From Attention Probing to Guidance Tuning for Multimodal Generation

Abstract

Unified vision-language models (VLMs) perform both text-to-image (T2I) and image-to-text (I2T) generation with one shared Transformer backbone, yet how the backbone processes the condition in either direction remains poorly understood. We study this question from the perspective of conditional guidance: we probe every attention head to show the activation difference between two branches of classifier-free guidance (CFG) with different conditional guidance. On two unified VLMs of autoregressive and masked-diffusion paradigms, the probing reveals a subset of heads that are more sensitive to the change of condition while others are agnostic to it. This characteristic pattern holds in both T2I and I2T generation and is largely stable across data samples, demonstrating the internal mechanism of unified VLMs. Motivated by this, we propose a new inference-time intervention approach for unified multimodal generation, named HeadTuned-CFG, which uses the probed condition-change awareness to select and tune the conditional guidance at the attention-head level. Experiments show consistent benefits in both multimodal generation directions of unified VLMs without additional training, and in particular substantial improvements over strong baselines in T2I generation. These results further support the key finding of our probing analysis: the observed pattern of functional specialization is an intrinsic property of unified VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.