acceptodds
Under review as a conference paper at ICLR 2027

Analyzing Preference Optimization with Synthetic Contrastive Reasoning Traces for Multi-Table Q&A

Abstract

Multi-table question answering requires models to retrieve relevant evidence, link schemas, and perform compositional reasoning across relational tables. Preference optimization over paired correct and incorrect reasoning traces is ideally suited for teaching LLMs such capabilities, but it is unclear which underlying mechanisms are most critical for this task: the training objective, or the properties of the contrastive reasoning traces. We study this question using synthetic contrastive reasoning traces for MMQA, consisting of validated positive traces and plausible negative traces generated by heterogeneous LLMs, and three open-weight LLMs (Qwen3-14B, Mistral-8B, and Llama-3.1-8B) evaluated on three table Q&A benchmarks, plus an out-of-distribution multi-table evaluation set built from BIRD. To investigate the role of optimization methods, we hold the preference pairs fixed and compare DPO, DPO with an NLL regularizer, and reference-free Contrastive Preference Optimization (CPO). We then further analyze the role of contrastive data diversity, quality, and source of the traces: whether heterogeneous generators for positive and negative traces strengthen the contrastive signal, how correct, faithful, and coherent the traces are under automated and human evaluation, and whether open-weight generators can replace proprietary ones. We find that CPO achieves absolute average improvements over Q&A supervised fine-tuning ranging from 10.7 to 19.3 percentage points, while DPO on the same pairs degrades two of the three models relative to positive-trace SFT and is only partially repaired by regularization. Removing the reference model improves every model over regularized DPO and is the larger of the two changes for all three; generating negatives with a different LLM than the positives gives a small but directionally consistent advantage; and an open-weight generator comes within 1–2 points of the proprietary one on every comparison. A case study and out-of-distribution transfer to three-table queries indicate that the negatives encode realistic multi-table failure modes, such as skipped constraint verification, that the trained models learn to reject.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.