Built to Read, Not to Recall: Language Skills over World Knowledge in Small Language Models
Abstract
Frontier-scale language models are commonly assumed to need their full scale. We argue instead that LLM abilities fall into two families with very different scaling behavior: world knowledge — stored facts and entities — improves steeply with parameter count, whereas language skills — comprehending, extracting, and structuring given text — are largely attainable at modest sizes when cultivated deliberately. Retrieval-augmented generation (RAG) and agent pipelines keep knowledge external and rely chiefly on the latter; a compact purpose-trained model can therefore replace frontier-scale components at a fraction of the cost and energy. We test this hypothesis with Meno-Lite-0.1, a publicly released bilingual Russian–English 7B-class model whose name inverts Plato's anamnesis: the knowledge lives in the corpus, the skill in the model — built to read, not to recall. Across a 0.5B–72B backbone size ladder, knowledge benchmarks climb steeply (closed-book quiz F1 grows 21× from 0.5B to 72B) while skill benchmarks plateau after 7B, and the memorization ratio — answering from parameters instead of the given context — rises with scale. The skill-trained 7B model sits above the entire ladder on knowledge-graph construction (0.468 vs. 0.416 harmonic mean against a 32B model), reaches 92% of the 32B level on open-book multi-hop QA while retaining its backbone's knowledge level, is the most context-faithful model measured, and serves Russian text 6.2× faster than the 32B engine. We further document protocol sensitivities that can masquerade as scaling behavior. Weights, training data, and code are publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.