Authorization-Conditioned Capability Isolation for Collaborative MLLM Inference
Abstract
Collaborative MLLM inference creates a mismatch between what a remote model can infer from a transmitted representation and what is authorized under the specific request. Existing privacy-oriented approaches primarily aim to reduce input reconstruction or semantic leakage, but do not explicitly control which capabilities a deployed model may exercise under request-level authorization. This work formulates authorization-conditioned capability isolation as transforming a shared visual representation so that a denied target contributes no additional capability beyond the model's No-Image behavior, while non-target capabilities remain close to their Raw behavior under the same transmitted state. We introduce an Authorization-Conditioned Bottleneck (ACB) that realizes this objective through policy-conditioned representation modulation, including token scaling, channel reweighting, rank-limited correction, and fixed-representation off-target preservation. In the primary evaluation on our constructed AuthCap, a 13-capability authorization benchmark, ACB achieves authorized utility of while reducing denied target capability to within of the No-Image reference. With off-target preservation, unrelated capability retention improves from to . These results demonstrate the feasibility of request-level capability isolation: suppressing unauthorized capabilities while substantially retaining unrelated visual capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.