Learning to Use What Models Know: Knowledge Retrospect Learning
Abstract
Large language models have made substantial progress on complex reasoning tasks, yet their performance remains less reliable when solving knowledge-intensive problems. In these settings, a model must not only access relevant facts, but also express them as intermediate premises during multi-step chain-of-thought reasoning. We study this gap through the lens of knowledge dispatch, focusing on whether accessible knowledge is actually invoked in the generated reasoning trajectory. To this end, we introduce a factorized diagnostic framework that separates isolated knowledge accessibility from in-chain knowledge use. Our analysis shows that knowledge points answered correctly in isolation are often not expressed during free-form reasoning. Further diagnostics indicate that these failures are more frequently associated with omission than incorrect invocation, and that the model can often use the omitted knowledge when it is explicitly provided. These findings suggest a knowledge dispatch bottleneck: relevant knowledge may be available to the model, but not reliably surfaced during generation. Motivated by this diagnosis, we propose Knowledge Retrospect Learning (KRL), an on-policy self-distillation framework that uses knowledge-augmented contexts during training to provide distributional supervision for problem-only reasoning. The student samples trajectories from the original problem context, while a knowledge-conditioned teacher retrospectively evaluates the same trajectories and transfers token-level signals to the student. Experiments on multiple knowledge-intensive reasoning benchmarks show that KRL improves reasoning accuracy, suggesting that training-time knowledge conditioning can help models surface relevant knowledge more reliably during generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.