acceptodds
Under review as a conference paper at ICLR 2027

From Estimation to Experience: Building Reusable Deployment Knowledge for LLM Parallel Training

Abstract

Training large language models (LLMs) in data center networks (DCNs) requires exploring a large space of parallelization and resource configurations. Yet real-system observations are costly to obtain and therefore sparse, making deployment experience scarce and slow to accumulate. A natural alternative is to use performance estimators for low-cost exploration, but existing estimators typically produce scenario-specific solutions rather than reusable deployment knowledge. We therefore ask whether low-cost estimation can be converted into reusable deployment knowledge, and whether accumulated experience can in turn guide subsequent estimation. To this end, we propose Est2Exp (Estimation to Experience), a unified framework that couples a performance estimator with an experience bank through bidirectional feedback. The estimator uses a hybrid-granularity directed acyclic graph (DAG) to capture computation-communication dependencies, and LLM-assisted calibration with limited real-system measurements to refine latency functions, including irregular behaviors that are difficult to capture with conventional physical models. The experience bank distills estimator outputs into reusable deployment knowledge, which is reused to warm-start subsequent exploration and jointly refined with the estimator as new real-system measurements become available. Experimental results show that Est2Exp achieves 99.3-99.9% DAG-level latency accuracy across evaluated configurations with limited real-system measurements. Reusing accumulated deployment experience further improves cross-scale search success from 73.3% to 98.3% while substantially reducing search overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.