Linear-Cost Unsupervised Elicitation from Pretrained Language Models
Abstract
Post-training language models depend on human annotators for labels, which becomes a bottleneck as models grow more capable: annotation becomes unreliable as tasks outpace human expertise. Unsupervised elicitation avoids that by drawing labels from the pretrained model itself. Internal Coherence Maximization (ICM) is one such method; it approaches golden-supervision accuracy across a range of labeling tasks. However, its simulated-annealing search over mutual predictability requires forward passes, hours for a few hundred examples, and per-dataset hyperparameter tuning. We show this search is largely unnecessary. Our method, Linear-Cost Unsupervised Elicitation (LC-UE), replaces it with two inexpensive phases. The first phase labels each group in order of the model's zero-shot confidence, conditioning only on the groups labeled before it, which optimizes an autoregressive form of coherence rather than mutual predictability itself. The second phase closes that gap with a few refinement sweeps, scoring candidate label flips under a cheap proxy against a frozen context and verifying each sweep with one exact computation. Both phases satisfy the task's consistency constraints by construction. Across TruthfulQA, GSM8K-verification, Alpaca, and two pluralistic alignment benchmarks, LC-UE matches or outperforms ICM and few-shot bootstrapping at cost for , with no per-dataset tuning. On TruthfulQA (), this amounts to an estimated fewer forward passes than ICM, and at least fewer even under a conservative bound. Taken together, LC-UE makes unsupervised elicitation practical at scale and across diverse tasks, without the computational cost or per-dataset tuning that has limited prior methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.