acceptodds
Under review as a conference paper at ICLR 2027

PolyGen: Fully Synthetic Vision-Language Training via Multi-Generator Ensembles

Abstract

Synthetic data offers a scalable solution for vision-language pre-training, yet current state-of-the-art methods typically rely on scaling up a single generative backbone, which introduces generator-specific spectral biases and limits feature diversity. We introduce **PolyGen**, a framework grounded in the *Generator-Invariance Hypothesis*: when images are rendered by an ensemble of architecturally distinct generators, generator-specific artifacts vary across sources while semantic content is shared by construction, and a model trained on the union is encouraged to encode the latter. We treat this hypothesis as a training principle with three testable implications, addressing ensemble composition, invariance regularization, and the timing of hard-negative introduction, each examined through dedicated ablations. PolyGen operationalizes it through a multi-positive strategy exposing the model to disjoint visual manifolds rendered from the same caption, together with a Programmatic Hard Negative curriculum enforcing fine-grained compositional understanding. By reallocating the same image budget from unique captions to multi-source variations, PolyGen yields a more robust feature space, outperforming the leading single-source baseline (SynthCLIP) by on aggregate multi-task benchmarks and on compositional reasoning, and, under an identical protocol, exceeding a real-data (CC3M) baseline on cross-modal retrieval while closing two thirds of the zero-shot gap.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.