A Framework and Checklist for Validity Threats in Foundation Model Research
Abstract
Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive. Instead, the community increasingly relies on research strategies that approximate the ideal experiment at a fraction of the cost: proxy experiments and scaling laws, observational studies with publicly available models, and single-run designs that leverage variation within individual training runs. In this work, we show that these research strategies perform causal inference in disguise. Specifically, they aim to measure treatment effects conditional on a specific training recipe. To estimate this treatment effect, efficient research strategies replace various parts of the ideal experiment with cheaper substitutes. Based on this analysis, we draw on concepts from the empirical social sciences and propose a unified framework to identify the implicit assumptions made by different research strategies. We find that each comes with a characteristic validity profile: proxy experiments gain statistical and internal validity at the expense of external and construct validity; observational studies face confounding and effect heterogeneity; and single-run designs are strained by interference between treated units. In case studies with recent papers, we find that these assumptions often receive insufficient attention. To overcome this problem, we propose a checklist that can help authors and reviewers identify threats to the validity of foundation model research designs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.