acceptodds
Under review as a conference paper at ICLR 2027

Learning to Solve Combinatorial Optimization Problems Turns LLMs into Multi-Paradigm Optimization Models

Abstract

Large language models (LLMs) are increasingly used for optimization through distinct paradigms, including optimization modeling, heuristic discovery, and direct text reasoning. However, these paradigms are typically developed and post-trained in isolation, leaving open whether capabilities learned under one paradigm can transfer to fundamentally different ways of problem-solving. We study this question through *cross-paradigm transfer*. We introduce *Reasoning-to-Code Optimization (R2CO)*, a source post-training setting where an LLM reasons about the structure of a combinatorial optimization problem, derives a solution strategy, and implements it as a self-contained executable program. To strengthen strategy-level learning, we further develop reflective reinforcement learning, which uses previously evaluated reasoning trajectories and verifier-derived diagnostic feedback to refine solution strategies. Despite being trained exclusively under R2CO, the resulting model consistently transfers to optimization modeling, heuristic discovery, and direct text reasoning tasks. Diagnostic evaluations further reveal improvements in constraint reasoning, planning, procedural reasoning, and multi-step mathematical reasoning. In contrast, alternative solver-grounded optimization post-training yields weaker transfer. These results suggest that cross-paradigm generalization is not an automatic consequence of optimization post-training, but depends on whether the training setting strengthens reusable reasoning capabilities that transcend a particular solution representation or interface, offering a new perspective on how LLMs can be trained for broader optimization competence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.