acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language Models as End-to-End Combinatorial Optimization Solvers

Abstract

Large language models (LLMs) have shown promise for automating combinatorial optimization (CO), yet existing approaches rely on textual representations of problem instances. Many real-world CO problems are naturally presented in visual forms, and text-based encodings may lead to spatial information loss and limit spatial reasoning capabilities. In this paper, we introduce a novel framework that empowers vision-language models (VLMs) to serve as end-to-end CO solvers by directly processing visual problem instances. We design visual-attributed instances that encode problem structures and basic features into images, and develop selfevolutionary visual prompt optimization to automatically discover effective visual designs. Our two-stage training strategy first uses supervised fine-tuning (SFT) to learn solution patterns from domain-specific solvers, followed by reinforcement learning to reduce constraint violations and improve solution quality. Additionally, we introduce inference-time visual augmentation that exploits structure-preserving transformations to generate diverse candidate solutions. Experiments on six NPhard CO problems show that our method achieves 100% feasibility, with optimality gaps ranging from 1.78% to 10.93%. It improves average feasibility over advanced reasoning models, including GPT-6-Astra and Gemini-2.5-Pro, by 26.5% and 44.3%, respectively, while also outperforming specialized heuristics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.