From Alignment Diagnosis to Joint Supervision in Multimodal Draft Training
Abstract
Multimodal draft training benefits from multimodal (MM) and text-only (TXT) supervision, yet the underlying alignment changes remain poorly understood. We study draft–target disagreement through an acceptance-related risk framework that decomposes distributional error into text-conditioned and visual-response components. Lower overall risk generally accompanies longer accepted sequences, while similar text-conditioned alignment can coexist with markedly different visual-response risk and acceptance. Controlled continuation training shows that high-risk groups outperform random and low-risk groups, and visual-response risk guides MM training more effectively than aggregate risk. However, useful within-modality rankings do not ensure effective cross-modal ranking: raw-score pooling overwhelmingly favors TXT examples in our candidate pools. We therefore compare risks relative to their modality-specific means rather than by absolute magnitude. Mean normalization preserves within-modality rankings and enables mixed continuation-set construction without predefined modality quotas. With an EAGLE draft for LLaVA-1.5-7B, mean-normalized joint construction improves macro-averaged accepted length across five benchmarks from 3.03 with uncalibrated raw-score pooling to 3.60, under the same 34K-example continuation budget and common warm-up.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.