A Unified Framework, Dataset, and Benchmark for Large AI Model Training and Inference Simulation
Abstract
Simulation of large AI model training and inference is essential for system design, performance prediction, and strategy optimization. However, existing studies and tools remain fragmented, and there is still no unified framework or standardized benchmark for systematic comparison. To address this limitation, we propose a unified framework for large AI model training and inference simulation that organizes existing methods from the perspective of the end-to-end execution pipeline and clarifies their key components, modeling targets, and applicability. Based on this framework, we collect runtime data from representative large AI models under diverse parallel configurations on a distributed computing cluster and publicly release an open dataset for large-model system simulation research at https://github.com/HongriJiujiu/LM-SimBench. We further establish a standardized benchmark for evaluating simulation methods in terms of accuracy, generalization, and efficiency. The proposed framework, dataset, and benchmark provide a standardized basis for analyzing and comparing simulation approaches, and facilitate more systematic research on large AI model training and inference simulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.