OptArchitect: Building Verifiable and Challenging Optimization Problems through Recursive Skill Composition
Abstract
Large language models are increasingly capable of solving existing optimization problems, making many established benchmarks less effective for evaluating and improving advanced reasoning. We present **OptArchitect**, a framework for automatically constructing challenging and verifiable optimization problems through recursive composition of reusable reasoning skills. Instead of directly prompting a language model to generate problems, OptArchitect first extracts structured skills from optimization papers, industrial documents, and solver cases, representing each skill through its problem semantics, mathematical structure, executable solver implementation, and transfer constraints. The framework then composes compatible skills through a structure-aware matchmaker and introduces explicit bridge operations to create nontrivial interactions between them. Generated problems are validated through executable optimization models, independent solver probes, structural consistency checks, and iterative skill refinement. This process supports recursive composition, enabling problems to be generated at increasing depths of structural complexity. Experiments show that current frontier reasoning models perform well on existing optimization benchmarks but exhibit substantial degradation on OptArchitect-generated problems. Across four reasoning models, average accuracy decreases from 86.1% on existing IndustryOR problems to 56.1% on recursively composed OptArchitect problems. The degradation becomes more pronounced as composition depth increases, with GPT-5.5 accuracy falling from 84.5% at depth to 33.3% at , while GLM-5.2 falls from 61.2% to 0%. The generated problems also exhibit substantially greater structural complexity, with the number of decision variables increasing from 40 at to 651 at , and the number of constraints increasing from 40 to 909. Ablation studies further show that structured skill representations, compatibility matching, bridge operations, executable solver verification, independent probing, and iterative refinement are all important for maintaining generation quality. Human evaluation confirms high problem correctness and modeling completeness, while cross-backbone experiments demonstrate that extracted skills transfer across different language models. Finally, models fine-tuned on OptArchitect-generated problems show improved performance on downstream optimization reasoning tasks. These results suggest that recursive, verifiable skill composition provides a scalable approach for constructing difficult reasoning data beyond the capabilities and limitations of existing optimization benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.