acceptodds
Under review as a conference paper at ICLR 2027

FoldUMM: Unlocking Efficient Inference in Unified Multimodal Models via Task-Specific Layer Folding

Abstract

Unified Multimodal Models (UMMs) couple image understanding and text-to-image generation within a shared Transformer backbone, but this shared design introduces substantial inference overhead. Through layer-wise redundancy profiling, we observe a structural divergence between the two tasks: understanding exhibits concentrated redundancy in deeper layers, whereas generation is more sensitive to depth reduction and lacks continuous redundant blocks. This task-dependent redundancy makes static layer pruning prone to cross-task degradation when transferred to UMMs. We propose FoldUMM, a task-aware layer-folding framework that realizes a "One Backbone, Multiple Paths" inference paradigm. FoldUMM formulates layer folding as a constrained shortest-path problem on a directed acyclic graph, and dynamic programming identifies task-specific folding paths under a target pruning budget. Lightweight Stitch LoRA modules are further inserted at fold points to repair representation mismatch. Experiments on Janus-Pro and OmniGen2 show that FoldUMM achieves a more favorable performance–efficiency trade-off, preserving substantially higher understanding and generation performance under structured pruning while providing practical inference speedup. Our code is available at https://anonymous.4open.science/r/FoldUMM/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.