From Decision Trees to Strong Reasoning Models for Tabular Data Classification
Abstract
Tabular data classification in high-stakes domains such as healthcare and finance requires both high accuracy and faithful, human-understandable reasoning. Symbolic decision trees provide explicit decision rules but offer limited semantic explanations, while adapting large language models (LLMs) to domain-specific tabular reasoning requires scalable and reliable reasoning supervision. We propose **ReSS**, a framework that combines symbolic decision structure with LLM semantic knowledge to curate reasoning data and train specialized tabular reasoning models. ReSS extracts decision paths from training instances whose tree predictions agree with their ground-truth labels. These paths, together with input features and labels, guide an LLM to construct natural-language rationales that follow the selected symbolic constraints and incorporate domain-specific interpretations. A scaffold-invariant augmentation strategy is proposed to expand the training dataset while preserving the symbolic constraints. The resulting data are used to fine-tune an open LLM that generates reasoning and predictions directly from input features, without requiring the decision tree at inference. We assess reasoning faithfulness through hallucination rate, explanation necessity, and explanation sufficiency. Experiments on medical and financial benchmarks demonstrate accuracy gains of up to \(10%\) over decision trees and standard fine-tuning baselines, alongside grounded and consistent explanations. Sensitivity analyses show that predictive accuracy remains relatively stable across the evaluated tree configurations, demonstrating that useful reasoning supervision can be obtained from trees of varying predictive quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.