Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models
Abstract
Hierarchical data is ubiquitous in the empirical sciences and most commonly analyzed with generalized linear mixed-effects models (GLMMs), a flexible hierarchical model class with interpretable parameters. Bayesian inference for GLMMs yields calibrated uncertainty estimates but requires Markov Chain Monte Carlo (MCMC); the gold-standard No-U-Turn Sampler (NUTS) is slow and must restart from scratch for every new dataset, model and prior specification. We introduce metabeta, a pretrained neural network for prior-amortized in-context Bayesian inference over GLMMs. Unlike previous neural posterior estimators that fix the prior at training time, metabeta accepts prior families and hyperparameters as direct inputs at test time, enabling zero-shot generalization to new prior specifications. Two set transformers and two conditional neural spline flows mirror the posterior's two-level structure (global parameters shared across groups, local parameters per group) and are trained jointly on millions of realistic simulated datasets with continuous, binary and count outcomes. By default, the flow posterior is refined by Independence Metropolis–Hastings against the unnormalized posterior; this yields tuning-free asymptotically correct inference two to three orders of magnitude faster than NUTS. Alternatively, metabeta warm-starts NUTS, yielding nearly identical inference 2–12× faster while removing over 98% of its divergent transitions. On controlled benchmarks with ground-truth parameters, metabeta matches NUTS in parameter recovery, posterior calibration and out-of-sample prediction; on real datasets with unknown parameters, priors and model structure, its posteriors closely agree with those of NUTS; and it remains faithful where amortized models typically struggle: under misspecified likelihoods, misspecified priors, out-of-distribution predictors, collinear designs and data-poor regimes. On real data, a full prior sensitivity analysis including model comparison takes a few seconds with metabeta against 7–17 hours for NUTS, tracing its posteriors throughout. Our model is open-source and open-weights and thus immediately deployable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.