acceptodds
Under review as a conference paper at ICLR 2027

Covering Language Model Outputs with Simplicial Complexes

Abstract

Language models (LMs) predict the next-token probability distribution over a large vocabulary, yet most of the probability mass is often concentrated on a small, context-dependent subset. This raises the question of whether LM outputs across contexts admit a compact representation. To answer this question, we first characterize this behavior of LM outputs using Bhattacharyya geometry on probability simplexes and simplicial complexes from algebraic topology. Our theorem connects probability mass discarded by truncation to our geometric dissimilarity measure. Because each output point lies near the low-dimensional simplex of its high-probability tokens, the output points are covered by a neighborhood of a simplicial complex, a union of simplexes. Next, we evaluate representations of LM outputs through the tradeoff among four metrics: . The first two measure the complexity of simplicial complexes, while the latter two assess approximation quality. To construct compact covers, we first consider collecting pointwise covers. For a single output, Top- and Top- truncation determine optimal simplexes in the tradeoff between and with fixed and . Collecting these simplexes yields simplicial complexes, which contain a prohibitively large number of simplexes in practice. To obtain a more compact complex, we formulate the construction task as a monotone submodular maximization problem and propose an algorithm for this problem. The algorithm provides the same worst-case guarantee as previous work. In the experimental section, we evaluate the construction using a natural language corpus (Dolma) and real-world LLMs (OLMo, Pythia). One of the resulting simplicial complexes uses 843 simplexes of dimension at most 4096 to cover 95% of the evaluated outputs while retaining at least 90% of each covered output's probability mass, providing a compact representation of approximately 50K-dimensional LM outputs over a 100M-token corpus.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.