acceptodds
Under review as a conference paper at ICLR 2027

CoRAL Synthesis: Learner-Relative Calibration of Data Complexity Composition for Improving VLM Generalization

Abstract

The data complexity of training examples can substantially affect optimization and generalization. Yet, data complexity remains an underspecified axis of synthetic data design, with complexity typically controlled using simple heuristics rather than principled criteria. This raises two questions: how data complexity affects downstream performance and how it should be leveraged to promote generalization. To study this in vision-language tasks, we construct hundreds of controlled VQA post-training settings spanning multiple complexity levels, using matched distinct subsets that isolate complexity from data scale, quality, and diversity. We find that complexity substantially affects out-of-distribution (OOD) generalization under a controlled setting, with some smaller subsets outperforming larger training sets. Moreover, our empirical analysis shows that appropriately mixing Easy, Medium, and Hard examples improves generalization over relying on any single complexity level, but realizing this benefit requires defining these levels relative to the target learner. We also find that existing complexity measures provide limited guidance for synthetic data curation. Building on these findings, we introduce CoRAL Synthesis, a scalable framework that operationalizes learner-relative complexity to govern synthetic data composition. CoRAL Synthesis consistently outperforms naive few-shot synthesis across data scales, with its 1K setting surpassing the vanilla 16K setting. Using CoRAL Synthesis, we release CoRAL-GeneralVQA and CoRAL-ReasoningVQA, two large-scale synthetic datasets for vision-language reasoning. Models trained on these datasets achieve strong OOD generalization; in particular, CoRAL-ReasoningVQA-4B outperforms the strongest evaluated 4B baseline by 5.30%p on average across 9 unseen vision-language reasoning benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.