Reading the Tea Leaves, Again: Can Large Language Models Predict Language Model Training Runs?
Abstract
LLM agents now propose research ideas and run the experiments that test them. However, they run every experiment to its end, whereas a researcher babysits a run, reading its configuration and the beginning of the curve, and stops the ones that show no promise. We ask whether LLMs can read a run the same way and build Tea Leaves to measure it, from 682 real training runs across pretraining, fine-tuning, and reinforcement learning (RL). A forecaster sees the first 20% of a run with its configuration, description, and code, and predicts the rest, which requires both curve extrapolation and a working knowledge of training dynamics. Across 11 LLMs and a range of forecasting baselines, the strongest LLMs match or exceed the best curve fitter on smooth loss curves (skill 0.66 against 0.62) and lead on noisy RL reward curves (0.43 against 0.13), but their skill falls as the runs get harder to read: on sparsely logged community runs they lose most of it, and no method reliably flags a failing run in advance. Our analyses identify what makes a run hard to read: its early curve is too noisy or too sparsely logged to extrapolate from. Knowledge of training dynamics helps where the curve misleads; related runs in the prompt cut the error on loss curves by up to 39%, and tools and finer input windows add little for the strongest readers. The benchmark and its database of runs will be released to test LLMs' understanding of training and help research agents decide which runs deserve their compute.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.