ParadigmCoT: Paradigm-Guided Chain-of-Thought in Large Language Models via Policy Optimization
Abstract
Chain-of-Thought (CoT) prompting may exhibit reasoning-path drift, logical inconsistency, and fabricated explanations, limiting the reliability of LLM reasoning. While structured prompting and process supervision improve local correctness and factuality, they provide limited holistic, principle-level guidance for the overall reasoning trajectory. Drawing on established instructive paradigms in both general and domain-specific fields, we propose ParadigmCoT: a general and flexible framework that injects instructive paradigms into the CoT process via policy optimization. At its core is an automated paradigm localization and evaluation module, which identifies paradigm-triggering steps and evaluates step-wise adherence throughout the reasoning chain. These paradigm alignment signals, combined with factuality and answer correctness, are incorporated into a fine-grained advantage adjustment in reinforcement learning to promote paradigm-aligned reasoning. ParadigmCoT is instantiated with general and medical paradigms, and evaluate it across diverse benchmarks and three LLM backbones. It achieves substantial gains over base models (up to 40.2% on general tasks and 34.5% on medical tasks), and outperforms RL-based baselines by 2.5%–3.4% and 5.8%–7.7% on average, respectively. This framework offers a new perspective on improving the quality and reliability of LLM-generated CoT, contributing to more reliable AI.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.