Books, Copyright, and the Future of Foundation Models: An Empirical Study of Fiction in Pretraining
Abstract
The use of books to train large language models has prompted numerous lawsuits against commercial AI providers, yet the reasons such data choices matter remain largely opaque outside industry labs. We ask a simple question: what role does fiction play during pretraining? To answer it, we pretrain 1B-parameter models from scratch under controlled token budgets on data mixes that include web text alone, web text with public-domain books, web text with contemporary fiction, and web text with both. Our results show that fiction substantially improves downstream model performance on creative writing tasks, with contemporary fiction yielding gains beyond those from public-domain corpora. We then examine what makes fictional text distinct from other pretraining sources, and examine whether synthetic data can recover its benefits. We find that models we trained on our synthetic data recipe achieve performance comparable to models trained on contemporary fiction, while making it harder to establish provenance through membership inference and extraction attacks. Together, these experiments offer the largest empirical study to date of copyrighted fiction in model pretraining, contributing evidence to debates over data governance, creators' rights, and how particular data domains shape AI capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.