Beyond OOD: Evaluating Large Foundation Models via Out-of-World Distribution
Abstract
Large Foundation Models (LFMs) have shown strong performance across diverse tasks, yet it remains unclear whether such performance reflects genuine cross-distribution generalization or interpolation over web-scale pretraining data. Conventional out-of-distribution (OOD) evaluation is ineffective, since the boundaries of LFM pretraining corpora are opaque and data contamination is difficult to rule out. Building on a world-model view of data generation, we make explicit a world variable that governs the input–output semantics and use its regular-world support to characterize the pretraining boundary. This leads to Out-of-World Distribution (OOWD), an OOD-like evaluation framework that first intervenes on this world variable and then evaluates cross-distribution transfer under the induced counter-world semantics. We instantiate OOWD across diverse generalization settings, including symbolic-to-verbal mathematical, program-to-natural-language, cross-style visual, and text-to-image transfer. Across these settings, we find that LFMs consistently transfer counter-world rules from the source domain to unseen target domains, whereas small models fail despite also fitting the source-domain data. Mechanistic analyses associate transfer with deeper layers and sparse causally relevant components. We further convert the principle into algorithmic design, OOWDAlign, for improving the generalization of scratch models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.