acceptodds
Under review as a conference paper at ICLR 2027

Spanning, Not Counting: Novelty in Language Generation in the Limit, from Finiteness to Compactness

Abstract

A generator can produce endlessly many distinct sentences and keep saying the same thing, and language generation in the limit, in the model of Kleinberg and Mullainathan, cannot tell the difference: it asks only that outputs be unseen. We ask that each output lie outside the span of everything said so far, for a span given by topics, a matroid, a consequence relation, or any monotone map that contains its argument. The right notion of a small set is then a finite core, a set spanned by finitely many of its own points: for the identity a finite set, for a finitary closure operator a set whose closure is compact in the lattice of closed sets. With finite cores in place of finite sets, the characterizations of Kleinberg and Mullainathan (NeurIPS 2024) and of Raman, Li and Tewari (COLT 2025) hold for every span, at every level when only the examples count as said, and uniformly when the generator's own outputs count too. The uniform price is free or fatal: as many examples as before, or no number at all, as for the paraphrases of one statement. At the other levels the generator's bets are the only obstruction, through a second mechanism that counting cannot see: poisoning, a bet valid for one language that together with examples of another spans it. Poisoning is impossible when the span is locally Noetherian, can be fatal beyond, as in a logic with explosion, and cannot be detected from inside single languages. Matroids, and several matroids at once, are the main instances.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.