Task-Free Visual-Token Repair for Frozen Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) rely on an internal interface that maps visual representations into the language model. When these representations are altered by quantization, compression, pruning, or other efficiency constraints, recovering downstream utility would ideally require optimizing them directly against the MLLM's downstream objective. However, this requires task-specific supervision and repeated forward/backward passes through a large MLLM, making end-to-end adaptation expensive and often impractical. We instead study task-free recovery at the visual-token interface. We introduce a lightweight repairer trained only from paired clean and corrupted visual representations, while keeping the visual encoder, multimodal projector, and language model frozen. Training requires neither task labels, prompts, language losses, nor LLM forward passes. Controlled perturbation experiments show that pointwise token fidelity is incomplete: representations with similar reconstruction error can induce substantially different downstream behavior, depending on their structure and direction. We evaluate this approach under practical corruption across Honeybee, LLaVA-v1.5-7B, and Qwen2.5-VL-7B. Under severe static 2-bit visual-token quantization, task-free repair recovers 81.8%, 72.5%, and 52.9% of the POPE degradation on Honeybee, LLaVA-v1.5, and Qwen2.5-VL, respectively, while objectives that preserve richer representation properties can outperform standard pointwise reconstruction. Learned image compression at approximately bpp provides a second, upstream corruption regime: repair recovers downstream utility across all three architectures, while robust and structure-aware objectives outperform MSE in matched settings. Overall, our results show that lightweight, task-free visual-token repair provides a practical alternative to end-to-end MLLM adaptation and helps identify which representation properties should be preserved under different forms of visual corruption.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.