acceptodds
Under review as a conference paper at ICLR 2027

Stochastic Estimation of Transduced Language Models

Abstract

Transduced language models (TLMs) compose a pretrained source language model with a functional finite-state transducer to induce a model over the target strings of the transducer. Computing the probability of a target prefix under a TLM amounts to summing the source-model probabilities of all strings that map to strings beginning with that target prefix. This set can be infinitely large, so prior work uses a computational shortcut and biased pruning to make the breadth-first search feasible. We instead propose an algorithm that samples source prefixes without replacement, reweighting each prefix by the inverse of its inclusion probability. Applying this correction gives an unbiased estimator of the TLM's prefix probability and lets us estimate the mass lost by pruning. The algorithm reduces the number of prefixes as more probability mass is added to the running estimate. This can save computation and guarantees that the run halts with probability one. The algorithm achieves a better compute–variance tradeoff on English text when converting a model over subwords to a model over whole words, and lower error on DNA-to-amino-acid transduction than other sequential Monte Carlo baselines. On the DNA-to-amino-acid transduction, it reduces runtime by orders of magnitude and makes scoring long target strings feasible. Repeating a published reading-time analysis substantially lowers estimated corpus surprisal while confirming its conclusions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.