Scaling Laws for Simulation Fidelity: Scale Alone Is Not Enough
Abstract
Both humans and AI can learn new skills by practicing with simulated human partners. Current LLM simulations can already provide meaningful learning signals, but they remain significantly less effective than human partners. Although LLM simulations look plausible on the surface, they often fail to reproduce human behavior patterns that are necessary for realistic and effective practice. We call this a gap in simulation fidelity. First, we introduce and validate distributional fidelity metrics, which we use to quantify fidelity gaps in two representative simulation tasks. We show that domain adaptation fails to close these gaps. This motivates our investigation of scaling laws for simulation fidelity. Web pre-training data is cheaper and more scalable than in-domain data, so we systematically study how simulation fidelity scales with pre-training compute on webtext. We discover that fidelity scales log-linearly with pre-training at smaller scales, but fidelity quickly asymptotes. Relative to other skill domains like creative writing and open-ended question answering, the rate of simulation performance improvement is much slower. Furthermore, domain adaptation does not alter the rate at which pre-training compute benefits model fidelity. In summary, we conclude that, even at scale, webtext data is an insufficient grounds for building more faithful LLM simulators. We recommend greater investment in larger-scale in-domain datasets than what are currently available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.