acceptodds
Under review as a conference paper at ICLR 2027

Unified Multimodal Models Do Not Learn in Unison: Active Probing of Cross-Capability Transfer

Abstract

Shared parameters in unified multimodal models do not imply shared learning. In our controlled setting, optimizing one capability can reduce the other by up to 3.70 percentage points on average, and similar passive gradient statistics can precede opposite receiver outcomes. We introduce Active Mechanism-Separating Probe (AMSP). AMSP first selects a bounded native low-rank probe that maximizes robust separation among action-distinct finite-response hypotheses. It measures source and receiver responses on a disposable branch, then commits a source-feasible intervention from the untouched optimizer state; insufficient evidence leads to abstention. On Show-o2-7B, AMSP reaches a joint understanding–generation harmonic score of 82.36%, exceeding Pareto LoRA by 1.68 points and equal-budget lookahead by 1.85 points; both margins remain positive under paired seed-and-combination resampling. On InternVL-U-4B, it again achieves the highest mean score at 84.42%. Lookahead repairs more destructive events, but AMSP preserves 98.15% of constructive updates versus 79.63% for lookahead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.