Generative Multi-modal Gradient Matching
Abstract
Multi-modal Large Language Models (MLLMs) rely on massive web-scraped corpora that are noisy, redundant, and weakly aligned with downstream objectives, leading to persistent failure modes such as object hallucination and brittle compositional generalization. Conventional text-to-image (T2I) synthesis offers a scalable alternative but remains passive: prompts drive sampling without feedback on training utility, producing images that are visually plausible yet redundant for fine-tuning. We present Generative Multi-modal Gradient Matching (GMGM), an active synthesis framework that, given a frozen T2I generator and a small real reference batch, optimizes per-prompt latent codes so that each generated image is simultaneously: (i) prompt-faithful, (ii) feature-aligned with real data via Sparse Autoencoder (SAE) distribution matching, and (iii) gradient-informative for downstream MLLM fine-tuning. Our key insight is that synthetic data should be optimized not merely for perceptual realism, but for alignment with the downstream model's learning dynamics. We prove that minimizing this tri-objective tightens the upper bound on the generalization gap, isolating covariate shift, feature coverage, and trajectory divergence. Empirically, relative to a strong prompt-matched synthetic baseline (TreeSynth), GMGM improves hallucination robustness (+5.6% POPE F1) and compositional reasoning (+58.1% Winoground), while matching its full-budget performance using only 10% of the synthetic data budget. Further experiments demonstrate GMGM's strong cross-architecture generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.