Filtering Low-Q Tokens: Improving Success Rates while Preserving Base-Model Successes
Abstract
LLMs are highly capable, yet they may still be imperfect with respect to a variety of target criteria, such as safety, contextual faithfulness, or reasoning. A natural goal is therefore to take a capable, trained LLM and augment it to increase its rate of successful responses. This augmentation should preserve the diverse successful outputs the base model already produces. We derive such an augmentation by testing each candidate token against the hypothesis that the final output will be successful. We show that the success probability of the full generation is lower-bounded by a quantity increasing in the power of this test. Our method, *Q-filtered decoding* (QFD), therefore follows from the Neyman-Pearson lemma, which identifies the optimal test of this form at a given false-rejection rate: a threshold on a token's probability of leading to a successful final output. At each decoding step, a small learned probe estimates this probability for every candidate token; those falling below the threshold are removed, and the remaining distribution is renormalized. Raising the threshold makes the filter intervene more aggressively, at the cost of altering more of the base model's successful outputs. Across mathematical reasoning, contextual faithfulness and harmful compliance, in seven task-model pairs, QFD significantly improves the success rate on six. We measure preservation by the drop rate, the probability that an intervention rejects a successful generation of the base model, and by the cosine similarity of successful completions to those of the base model. No baseline, including LoRA fine-tuning and activation steering, dominates QFD: at 28 of its 35 settings, no baseline reaches as low a drop rate, and at the rest, those that do are less successful. On four of the six pairs, QFD's largest improvement in success rate is comparable to the best baseline's, yet it drops 54-87% of the base model's successful generations instead of 96-100%, and its completions are closer to the base model's in cosine similarity. QFD slows generation only by 1.2% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.