acceptodds
Under review as a conference paper at ICLR 2027

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

Abstract

Direct Preference Optimization (DPO) has emerged as an efficient alternative to PPO-based RLHF, yet its application to knowledge-intensive generation tasks reveals a fundamental limitation: standard preference signals—whether from human annotators or LLM judges—exhibit a systematic verbosity bias that rewards linguistic fluency over logical correctness. We show empirically that this reward signal blindspot leaves a large logical alignment gap: SFT-trained models achieve NLI entailment scores of only 0.05–0.22, despite producing fluent, confident-sounding text. We propose RLearner-LLM, a framework that resolves this through Hybrid-DPO: an automated preference construction pipeline that fuses a DeBERTa-v3 NLI entailment signal with a verifier LLM score, eliminating the need for human annotation while overcoming the "alignment tax" that degrades fluency under single-signal optimization. Evaluated across five academic domains (Biology, Medicine, Law) with three base architectures (LLaMA-2-13B, Qwen3-8B, and Gemma 4 E4B-it), RLearner-LLM achieves up to 6x NLI improvement over SFT baselines, with NLI gains in 11 of 15 (architecture, domain) cells and consistent answer-coverage gains. On Gemma 4 E4B-it—a 4.5B-effective-parameter dense model—Hybrid-DPO lifts NLI entailment in four of five domains (from +11.9% on Auckland Law to +2.4x on UK Medicine Year 2), with faster inference across all five, demonstrating that the approach scales down to compact modern base models without losing the alignment-tax mitigation. Our Qwen3-8B RLearner-LLM wins 95% of blind pairwise comparisons against its own SFT baseline, while GPT-4o-mini in turn wins 95% against our concise RLearner-LLM output—a result that, viewed alongside the 69% win rate the same judge awards a verbose SFT baseline over our DPO-aligned model, reproduces the verbosity bias we document on a frontier comparator and reinforces the case for logic-aware automatic metrics (NLI, ACR) over LLM-as-a-judge on knowledge-intensive generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.