Different Learning Trajectories, Same Conclusion: An Investigation of LLMs via Differential Warrant Learning
Abstract
Reasoning traces are increasingly used as training supervision, yet reaching the same conclusion does not imply learning from the same inferential support. We ask whether changing the validity of a critical warrant, the inferential rule that licenses a targeted inference, changes a model’s reasoning behavior and how the resulting learning difference is organized internally when the reasoning instance, target inference, and training exposure are held fixed. To answer this question, we introduce Differential Warrant Learning (DWL), a matched counterfactual training framework that isolates this contrast. Across experiments on four different LLMs, Healthy (valid-warrant) training consistently yields a higher reasoning-sensitive Terminal validity margin than matched Severe (invalid-warrant) training, although the effect size varies substantially across models. Bidirectional activation transplantation further identifies later, prediction-associated representations that causally carry the Healthy–Severe behavioral contrast across all four settings. How this learned difference matures through depth, however, is model-dependent: representational magnitude does not track causal influence on behavior, and the Healthy–Severe representational difference decreases even as its decision alignment and finite causal leverage increase in one model. Across three model settings, independently constructed validity-related directions—built without Healthy/Severe condition labels or the Terminal-margin gradient—recover the predicted model-specific depth ordering and causally steer Terminal-validity behavior. Finally, activation-level causal expression does not determine where Healthy parameters can be productively inserted into the Severe-trained model. Together, these results distinguish representational difference, causal expression, and parameter replaceability as related but non-equivalent properties of the same matched learning contrast.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.