acceptodds
Under review as a conference paper at ICLR 2027

From Algorithms to Reasoning: Operations Research as Process Supervision for Language Models

Abstract

Autonomous combinatorial optimization remains challenging for language models, while operations research provides well-established algorithms for these problems, and internalizing these algorithms as learned solution strategies could enhance the models’ optimization capabilities. We propose a training framework that uses operations research as verifiable process supervision, translating algorithmic structures and mathematical evaluation criteria into learning signals for intermediate decisions. The model constructs and revises candidate solutions through natural-language reasoning, while a verifier used during training evaluates and compares alternative decisions from the same state using constraint satisfaction, objective values, and algorithm-specific quality criteria. Under shared process-quality criteria, the framework combines local preference learning weighted by quality gaps with reinforcement learning over complete trajectories to jointly train local decision-making and global coordination. Experiments across six classes of combinatorial optimization problems demonstrate improvements in feasibility rates and solution quality. This approach to algorithm internalization could strengthen language models’ reasoning over long-horizon tasks in real-world settings, providing a methodological foundation for agent decision-making.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.