What's in a Tree ? Hierarchical -gram models as controlled approximators of language
Abstract
Hierarchical models have been proposed as synthetic settings for studying language. However, the relevance to realistic language corpora is under-explored. In this work, we introduce tree-grams, autoregressive generative -gram models that model the conditional probability tensor through a hierarchy of order-2 and order-3 merges augmented with normalization layers, residual connections, and MLPs. At fixed , a rank hyper-parameter, , controls model capacity, yielding an interpretable and systematically scalable family of approximators of language. Tree-grams are shown to out-perform existing -gram baselines for higher , while remaining competitive with deep neural sequence models. We demonstrate the utility of such controlled approximations by using tree-grams as generative models. In particular, we show that hyper-parameter transfer under P can be made more robust when narrower models are tuned on simpler data. In addition, tree-grams allow us to investigate the effect of pretraining and finetuning distributions. We show that performance increasingly worsens as the complexity of the finetuning distribution increases. However, the pre-training distribution complexity has little effect in a capacity-limited regime, while there is a non-monotonic effect in a horizon-limited regime. Tree-grams are interpretable by construction. In particular, terminal nodes in different layers have receptive fields whose information can be independently recovered through linear probes and activation patching. When allowed to route across candidate topologies, we observe emergent sparsity where tree-grams produce concentrated path probabilities fit by a power law with a generalized exponential cutoff.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.