acceptodds
Under review as a conference paper at ICLR 2027

UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward

Abstract

While Large Language Models (LLMs) have demonstrated significant potential in natural language processing, complex general-purpose reasoning—requiring multi-step logic, planning, and verification—remains a critical bottleneck. Although Reinforcement Learning with Verifiable Rewards (RLVR) has succeeded in specific domains, the field lacks large-scale, high-quality, and difficulty-calibrated data for general reasoning. To address this, we propose UltraLogic, a framework that decouples the logical core of a problem from its natural language expression through a Code-based Solving methodology to automate high-quality data production. The framework comprises hundreds of unique task types and an automated calibration pipeline across ten difficulty levels, which is dynamically adjustable to remain effective and future-proof as base model capabilities evolve. Through feedback-guided task refinement and difficulty recalibration, it supports a data-centric approach to recursive self-improvement (RSI), producing verifiable training data matched to evolving model capabilities. Furthermore, to mitigate binary reward sparsity and the Non-negative Reward Trap, we introduce the Bipolar Float Reward (BFR) mechanism, utilizing graded penalties to effectively distinguish perfect responses from those with logical flaws. Our experiments demonstrate that BFR, combined with a difficulty matching strategy, significantly improves training efficiency, guiding models toward consistent logical policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.