TASK–DRIFT LEDGER: PRE-HOC CALIBRATION OF MODULE-SPECIFIC ANTI-FORGETTING STRENGTHS FOR MULTIMODAL LLM ADAPTATION
Abstract
Fine-tuning multimodal large language models (MLLMs) can erode capabilities acquired before adaptation. Parameter-space methods constrain model updates to limit forgetting, yet their intervention strengths are usually inherited from defaults or chosen through costly sweeps. The appropriate strength varies across tasks and between the cross-modal projector and language model; constraining one module can also shift adaptation pressure to the other. We introduce Task–Drift Ledger, which selects a strength for each module before full adaptation. A rank-4 variational probe fitted for 1,000 steps on 2,000 downstream examples estimates the task-induced parameter response. Combining this response with downstream sensitivity and reusable upstream sensitivity, the ledger scores the surrogate costs of retaining and suppressing the response. It sets module strengths while preserving each underlying method’s parameter-selection rule. On LLaVA-1.5-7B, the selected configurations improve the within-task score over operator references in all nine settings, and the strengths improve all 11 evaluable higher-rank configurations. With controller constants unchanged, the method also improves the balance between forgetting and downstream performance in most evaluated Qwen2.5-VL-7B captioning settings. In our pipeline, one probe costs about 0.5 GPU-hours, compared with 26.6–79.8 GPU-hours for a training-time strength sweep. Conditional sweeps further show that the selected strengths lie in useful regions close to the best measured points at substantially lower calibration cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.