acceptodds
Under review as a conference paper at ICLR 2027

Beyond Laplace: How Mamba In-Context Learns Constrained Switching Markov Chains — Estimation Gaps Close, Statistical Limits Do Not

Abstract

Recent work showed that Mamba exactly represents, and learns in training, the Laplacian smoothing estimator for unconstrained Markov chains. We study switching Markov chains whose transition kernels share one stationary distribution and whose switches are hidden from the predictor. This setting poses two distinct problems, estimating the constrained kernel and identifying the kernel currently in effect. For estimation, the stationarity constraint couples the kernel rows in the posterior, so Laplacian smoothing is strictly suboptimal for any number of states. We show that the difference between the constrained Bayes predictor and Laplacian smoothing lies in the readout rather than in the state. Readouts that use only the current row's counts cannot reach this predictor, whereas a log-linear softmax readout on the same count state represents it exactly for two states, with a width that grows linearly in the sequence length and a lower bound of the same order. Probes show that trained Mamba stores the transition counts linearly in its hidden state and combines counts across rows at the readout, closing most of the estimation gap. For identification, even the optimal filter with infinite memory incurs a strictly positive residual loss for any number of states and any order, and a Wald-type lower bound shows that switches cannot be detected instantly. Trained Mamba and Transformer models approach this statistical limit but do not exceed it. Provable unpredictability against well-trained sequence models must therefore rest on the statistical limit created by hidden switching rather than on the difficulty of estimation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.