acceptodds
Under review as a conference paper at ICLR 2027

Exactly Verifiable Reinforcement Learning for LLM-Written Rules

Abstract

An NP-hard problem is intrinsically difficult to solve optimally. In practice, one has to resort to heuristic algorithms, that is, hand-coded rules that provide quick decision-making and solutions of adequate quality. Workload scheduling is an example. In a large computing system, more jobs arrive than can be served. Thus, there is a need to create rules for selecting the next job to serve and to modify them as necessary due to changes in workload and/or goals. A rule can be expressed as a short program, hence a language model can write one, and an exact simulator can judge it by re-running a recorded log of real requests under that rule. The latter creates a verifiable reward. We present , a framework which approaches NP-hard problems by learning their heuristics. It trains a language model that writes rules and a selector that chooses among them. The rule writer proposes new rules to deal with shortcomings of the current set, the simulator verifies the candidates, and only valuable rules are selected. At runtime, the selector picks new rules to test and makes the scheduling decision. Using the BurstGPT public trace of inference requests, with the Qwen3.5/3.6, Nemotron 3 Nano, and Nemotron 3.5 Lightning model family, RuleForge manages to match the accuracy of a similarly sized search on the same set of rules with one-sixth the number of replays, meet goals not met by the search, and get within 3.3 points of the optimum of the complete trace in under one-twentieth of the number of replays. On the benchmark job-shop problems, the learned selector performs much better in meeting scheduling goals, by 7.2 to 10.0 points, and matches the schedule quality of an exact solver with one-third the computation time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.