acceptodds
Under review as a conference paper at ICLR 2027

Post-Training Open Multimodal Models for Optimization Modeling from Text and Images

Abstract

Large language models (LLMs) have advanced automated optimization modeling, but existing methods largely assume that problems are fully specified in text. In practice, text and images often provide complementary information needed to formulate a problem. While existing multimodal optimization modeling benchmarks reveal a substantial performance gap between open and proprietary multimodal large language models (MLLMs), they do not provide training data or post-training methods for this task. We therefore ask whether task-specific post-training can equip a compact open MLLM with this capability. We introduce MLLM4OR-Corpus, comprising 22,107 solver-verified instances. We partition the corpus into disjoint training, development, and test subsets for post-training, model selection, and in-domain evaluation, respectively. For external evaluation, we use MM-OptBench and introduce LLM4OR-VL to cover a wider range of problem sources and description styles. LLM4OR-VL comprises 878 multimodal problems reconstructed from eight text-only optimization benchmarks. Using these resources, we introduce MLLM4OR, an open 9B-scale multimodal LLM trained to generate mathematical formulations and corresponding solver code from multimodal specifications. To this end, we first use supervised fine-tuning (SFT) to connect text-image evidence to formulations and solver code, comparing training answers with no explanations, concise explanations, or detailed explanations of the modeling and implementation steps. We then propose a solver-guided preference construction method for direct preference optimization (DPO), contrasting verified answers with the model's own execution and modeling failures. Extensive experiments show that including concise explanations for some training problems and omitting them for others yields the strongest overall SFT performance. DPO further improves overall accuracy and performance on MM-OptBench and LLM4OR-VL. These findings support task-specific post-training for automated multimodal optimization modeling with compact open MLLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.