Ikigai is All You Need: Non-Gradient Referent-Ledger Construction Escapes the Bounded-Window Ceiling
Abstract
Gradient-trained language models pay a fixed, growing compute cost to exploit more data: acquiring new knowledge means reprocessing the entire corpus at full training cost. A constructive text generator built with no neural network, no gradient descent, and no transformer does not share this problem. Its acquisition cost grows sub-linearly with corpus size (wall-clock time N^0.51 across a 122x corpus sweep, process memory flat at about 50 MB for its language model), and, reversing every gradient-scaling law we are aware of, its per-sentence generation speed and reliability actually improve as the underlying corpus grows, by 3.4-4.8x across independently replicated seeds, simply because richer real evidence makes its constraint-satisfaction search easier rather than harder. The mechanism also closes a structural gap that has limited non-gradient generation: any admissibility predicate evaluated over a bounded context window is provably capped at a regular language (Theorem A), and every surface-marker escape mechanism we constructed and tested needed to reintroduce a fixed convention just to stay decodable (Finding B). Unbounded, non-surface generator memory, an exact-identity referent ledger, escapes both, with structural complexity that keeps growing with the tracked-referent pool rather than saturating. At equal training-corpus size, two independent blind human raters judged this mechanism more natural than a converged gradient-trained transformer baseline on the exact matched corpus at matched diversity (2.31 vs 1.94 of 3, p=0.011), and it matches that baseline on two standard diversity and copying measures under the same matched-corpus condition (Distinct-2: 0.766 vs 0.773; copied 7-grams: 0% vs 0.7%) — all with zero backpropagation, zero GPU requirement, and flat RAM regardless of corpus size for its constant-work language model. One axis it does not yet win is perplexity under an external pretrained model, which by construction rewards statistically generic text; we report it rather than omit it. The same clustering mechanism can also directly organize a pretrained transformer's own embedding weights, the literal parameters rather than a distillation, into categories that tie or exceed our from-scratch results. It does not yet match transformer-scale generation unconditionally, and we present it as a first, concrete step in an active research program toward that goal rather than a finished system.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.