acceptodds
Under review as a conference paper at ICLR 2027

Lead into Gold: The Surprising Effect of Symbolically Generated Deduction Data on LLMs

Abstract

Synthetic data now supplies much of the instruction tuning of language models, and it is usually written by another language model with the goal of sounding as natural as possible. We ask whether data at the opposite extreme also helps, a corpus of multi-hop deduction problems and their solutions that a short program generates with no model in the loop. The documents are correct by construction, but they carry no world knowledge and they name their constants abstractly. We replace a small fraction of a standard instruction-tuning mixture with such documents while holding the number of training examples fixed. On Qwen2.5-7B, replacing 5 to 10 percent of the mixture raises accuracy by 14 points on ProofWriter, 6 points on FOLIO, 7 points on the multi-step subtasks of BIG-Bench Hard and up to 7 points on the quantitative items of GPQA-Diamond, and it leaves general benchmarks unchanged. These benchmarks cover formal logic, language inference, word problems and the sciences, so the gain is not confined to the domain of the added documents. The gain on ProofWriter holds on models of 3 families and from 3B to 32B parameters. As little as 1 percent of the mixture teaches the model to write long derivations that a symbolic checker accepts, an ability that no baseline model in our study has. Two mechanisms account for the gains. The model learns to carry an intermediate conclusion forward through a chain of steps instead of restarting from the premises, and it stops preferring one answer label over the others, a bias we trace to the instruction corpus itself. Data of this kind costs almost nothing to produce and needs no teacher model, so it is available to anyone who trains for multi-step reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.