acceptodds
Under review as a conference paper at ICLR 2027

How Many Dimensions Do You Actually Need? Testing Embedding-Dimension Rules on Their Slope, Not Their Value

Abstract

Rules of the form , where counts the items to be represented, are widely used to choose an embedding dimension. They are usually validated by evaluating the rule at one and comparing it with a conventional width, a check that a wide family of constants and functional forms passes equally well. In this work, we measure how the required dimension actually moves with , holding the task fixed and varying over up to decades in associative recall with an attention head, identity retrieval and topic recovery in word embeddings, and node classification on citation graphs. We find that the dimension a task needs is set by what it must distinguish rather than by . On the same synthetic embeddings with few planted topics, recovering the topics needs as many dimensions as there are topics whatever is, while telling the words apart needs a width that grows with . We also find that identity tasks on different data grow at rates that differ by more than an order of magnitude, and that classification tasks show no detectable common growth. Along the way, we show that dimension floors estimated relative to a noisy curve's maximum acquire a spurious slope in , and give an estimator that does not. We hope the protocol lets future dimension rules be stated for a task and tested on their slope.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.