Cues as Compasses: Cross-Modal Condition Steering for Unified Multimodal Models
Abstract
Unified multimodal models perform image understanding and generation within one architecture, yet architectural unification does not guarantee that each task uses its cross-modal condition effectively. Analyzing attention at the decisive stage of each task, we find a common conditioning gap in two task-specific forms: before answering, the model revisits the image but spreads its attention diffusely, whereas during generation, attention to the instruction decays early, leaving later updates weakly constrained. We propose Cross-Modal Condition Steering (CMCS), a training-free framework that contrasts each condition with a targeted reference to adaptively decide where, how strongly, and when to steer according to task demands. For understanding, CMCS redirects the model's existing visual attention toward instruction-relevant regions, and accepts the resulting answer change only to the extent that the reference evidence supports it. For generation, it reinjects, for every semantic unit, the residual between predictions under the original prompt and a counterfactual that alters that unit, shaped spatially at a fixed magnitude and scheduled by the unit's stage-dependent prior. Without updating any parameters, CMCS dynamically adapts to the current query and answer without degrading the original model performance, improving accuracy on visual understanding benchmarks and strengthening prompt adherence and compositional alignment on image generation benchmarks for both backbones, with further improvements in instruction adherence for image editing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.