Thinking inside the brackets: Grammar-Constrained Decoding
Abstract
LLM reasoning is a core driver of model performance, but it is costly. Practitioners typically use prompting and fine-tuning as axes of control to reduce the thinking-verbosity of reasoning models. We identify and exhibit grammar-constrained decoding (GCD) as a third axis of control. We structure reasoning traces by a small context-free grammar at decode-time (e.g. open a subgoal, derive, add a fact, check, close) which is compatible with free-text payloads. On GPQA-Diamond and AIME24, GCD on Qwen3.6-27B reduces median generated tokens by 98% and 94%, while holding or improving accuracy over chain of thought. While GCD reliably makes reasoning succinct, we further identify and characterise cases where controlling reasoning is detrimental, or ineffective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.