Opaque Multimodal-LLM Coordinated Unlearning
Abstract
Multimodal large language models (MLLMs) enhance LLMs' capabilities but raise significant privacy and compliance concerns, creating a need for machine unlearning (MU). Existing MU methods for MLLMs predominantly rely on white-box parameter updates, which are impractical in Large-Model-as-a-Service (LMaaS) settings with restricted model access. Current API-level alternatives adopt unimodal interventions that fail to address the entangled multimodal representations in MLLMs, leading to ineffective cross-modal forgetting. To overcome this, we propose OpaqueMU, a multimodal unlearning framework under a split-access LMaaS setting, where the Transformer backbone remains opaque while input-side representations and output scores are accessible. OpaqueMU introduces a coordinated two-stage architecture: a Manifold Sanitization (MS) module that selectively disentangles sensitive information within modality-specific representations prior to fusion, and a Distribution Alignment (DA) module that suppresses residual forget-related signals at the output level while preserving retained knowledge. Experiments across two multimodal unlearning benchmarks demonstrate that OpaqueMU achieves a favorable forgetting–utility trade-off while robustly preserving retained knowledge, alongside a 66% reduction in inference latency compared with the strongest baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.