FoldUMM: Unlocking Efficient Inference in Unified Multimodal Models via Task-Specific Layer Folding
Abstract
Unified Multimodal Models (UMMs) couple image understanding and text-to-image generation within a shared Transformer backbone, but this shared design introduces substantial inference overhead. Through layer-wise redundancy profiling, we observe a structural divergence between the two tasks: understanding exhibits concentrated redundancy in deeper layers, whereas generation is more sensitive to depth reduction and lacks continuous redundant blocks. This task-dependent redundancy makes static layer pruning prone to cross-task degradation when transferred to UMMs. We propose FoldUMM, a task-aware layer-folding framework that realizes a "One Backbone, Multiple Paths" inference paradigm. FoldUMM formulates layer folding as a constrained shortest-path problem on a directed acyclic graph, and dynamic programming identifies task-specific folding paths under a target pruning budget. Lightweight Stitch LoRA modules are further inserted at fold points to repair representation mismatch. Experiments on Janus-Pro and OmniGen2 show that FoldUMM achieves a more favorable performance–efficiency trade-off, preserving substantially higher understanding and generation performance under structured pruning while providing practical inference speedup. Our code is available at https://anonymous.4open.science/r/FoldUMM/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.