Mimir: A Sub-2B Hierarchical Reasoning Model Trained on Permissible Data Punches Above Its Weight Class
Abstract
Training capable language models typically relies on web-scale pretraining corpora followed by multi-stage post-training, creating substantial computational, as well as data-governance barriers for settings in which training data must satisfy explicit permissibility constraints. We introduce Mimir, a 1.78B-parameter language model based on the Hierarchical Reasoning Model Text architecture, a variant of recurrent transformers where the layer stack is split into L and H stacks that are iterated and interleaved. Mimir is trained from random initialization directly on a curated mixture of instruction, reasoning, tool-use, multilingual, and synthetic data. The training mixture comprises 161 datasets and 70.5B sampled tokens per epoch for Mimir v1, with substantial coverage of both English and Danish and with synthetic replacements for data sources that do not satisfy the target permissibility criteria. Across 20 benchmarks, Mimir substantially outperforms all sub-2B reference models, demonstrating that our permissible data approach produces state-of-the-art results for both high-resource languages (English, 1B speakers) and low-resource languages (Danish, 6M speakers) and is competitive on specialized domains such as mathematical reasoning and code. We conduct two memorisation audits, showing that memorisation risk is generally low even though the model is trained for multiple epochs. We also perform inference time ablations on recurrent depth, demonstrating graceful degradation for recurrence changes of the L stack. Our results provide evidence that hierarchical recurrent language models can learn effectively from post-training-oriented, permissibility-constrained data mixtures without relying on conventional web-scale pretraining. Mimir outclasses the original HRM-Text 1B, as well as the looped transformer Ouro, and represents the new state of the art among frontier sub-2B models, and competes with larger frontier models such as Gemma E4B and Qwen 3.5 9B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.