acceptodds
Under review as a conference paper at ICLR 2027

Attributing Representational Geometry to Data, Architecture, and Training

Abstract

Any observable property of a large language model (LLM) stems from three sources: the data it reads, the model architecture (the network and tokenizer), and the training process. We argue that training produces little of the representational structure we can measure in an LLM. Instead, most of that structure comes from the data and the architecture. Training adds one thing: the model's ability to tell what a text says from how it is written, and this ability appears early in training. Existing evaluation methods such as scaling laws and ablations compare trained models with one another and do not separate the contributions of these three sources, so claims about what a model has learned, and comparisons between models based on their representations, may rest on properties that training did not produce. To identify the contributions of these three sources, we propose DATS (Data–Architecture–Training Separation), a framework that measures the same documents in three ways and attributes each property to its source: a property already shown by bag-of-words (BoW) representations is attributed to the data, a property first shown by untrained networks to the architecture, and a property shown only by trained models to training. Applied to the geometry of document representations across four model families, DATS finds that (1) the ordering of document types by intrinsic dimension is already present in the data; (2) the similarity between two ways of writing the same text is already present in the word counts, and the tokenizer decides whether an untrained network keeps it; (3) only trained models place two ways of writing the same text together, and in Pythia they do so within the first few percent of training, while the shape of the representation keeps changing to the end.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.