acceptodds
Under review as a conference paper at ICLR 2027

Before Training, Which Multimodal Data Will Help? Resolution Limits Across Granularities and Learners

Abstract

Choosing a multimodal training corpus poses a catch: we learn whether it helps only after paying for the training run. When there are more corpora than we can try, can frozen representations tell us which ones are worth that cost? We test this question on three paired audio and text datasets for sentiment and emotion. Our sealed evaluation fixes each decision before candidate utility is revealed and repeats the comparison with four downstream learners. The answer depends on selection granularity. With logistic regression, centroid distance chooses the oracle source in 44 of 45 decisions, yet fails to rank buckets from that source. The separation between sources is 22.3 times the typical within-source spread; matching bucket size and class balance does not recover the ranking. Utility also changes with the learner and fusion rule, while candidate-specific versions of LEEP and LogME remain unreliable as direct bucket rankers. We use these signals to form a shortlist rather than replace downstream evaluation. Before opening a new repartition family, we fix a procedure that routes by distance, ranks by adapted LEEP, evaluates the top two candidates with the full downstream recipe, and keeps the target-only model if neither helps. Across three targets, it improves validation macro F1 by 0.01084 on average, with a 95% hierarchical bootstrap interval of , while avoiding 87.5% of full candidate runs. The result locates a practical boundary: prospective scores can guide a training budget, but they are not universal measures of dataset quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.