Beyond a Single Surrogate: Consistency-Weighted Alignment for Instruction-Free Multimodal Extension
Abstract
Extending a frozen LLM to a new modality can be reduced to training a lightweight projector, provided the alignment target is the LLM's own response distribution rather than external annotation; recent instruction-free alignment-only methods realize this with a caption as a semantic surrogate, whose supervision captures semantics well but — as prior work itself acknowledges — may under-represent super-semantic attributes. We take that acknowledged limitation as our starting point: surrogate fidelity is currently neither measured per sample nor acted upon, so when a single caption fails to stand in for the signal, the projector is silently taught to reproduce content the encoder never carried. **CoSA** makes surrogate fidelity an explicit, per-sample quantity: it constructs multiple surrogates per sample at differing granularity and viewpoint — deliberately spanning the super-semantic attributes a single caption tends to omit — elicits a response from the same frozen LLM for each, and uses the disagreement among these responses as a fidelity estimate, weighting rather than filtering samples so low-fidelity examples contribute proportionally instead of injecting unbacked supervision. It converts the previously implicit anti-shortcut argument into an **explicit constraint plus a measurable diagnostic** — a lightweight reverse probe reads surrogate-relevant attributes back out of the projected representation, while a modality-ablated control set quantifies how much accuracy survives when the signal is removed — and exploits the fidelity estimate for surrogate-aware active sampling, directing the generation budget where surrogate coverage is weakest. On MMAU-Pro and four audio-language benchmarks, CoSA improves over single-surrogate instruction-free alignment by 2.6 points while keeping the encoder and LLM fully frozen and using 42% of the alignment data; we find the previously reported saturation is not purely an encoder-capacity ceiling — over a third of the gap on closed-set tasks is recovered by reweighting alone — and the reverse probe reduces language-prior shortcut behavior, measured as a 4.8-point accuracy drop on audio-absent control inputs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.