RELIVE-OR: Retrospective Learning through Integrated Validation of Experience for Operation Research Modelling
Abstract
Synthetic data is increasingly used to fine-tune large language models (LLMs) for operations research (OR) modeling, yet scaling synthesis alone does not necessarily produce better OR modelers. We view each problem–program pair as a modeling experience and OR training as a process of acquiring, validating, revisiting, and reusing such experiences. This perspective exposes a key limitation of existing synthesis pipelines: beyond semantic, code-structural, and outcome consistency, a synthetic experience should remain consistent with the OR model specification represented by its trusted seed. Otherwise, it may remain executable and numerically valid while becoming misspecified with respect to the intended modeling principle. We propose RELIVE-OR—REtrospective Learning through Integrated Validation of Experience for OR modeling—with the objective of maximizing structural freedom while preserving OR model specification. Starting from a small trusted seed set, RELIVE-OR synthesizes structurally diverse experiences directly in executable code space and applies closed-loop verification and repair for semantic, code-structural, and outcome consistency. The verified experiences are first used for broad consolidation to expose the learner to diverse OR structures. RELIVE-OR then retrospectively estimates the influence of each verified experience on a trusted seed objective and re-selects experiences that remain aligned with the intended OR model specification for focused consolidation. Finally, the verified experience bank is reused at test time as a non-parametric reference memory for gradient-guided selection among executable candidates, without additional parameter updates. Across six established OR modeling benchmarks, RELIVE-OR enables an 8B model to achieve state-of-the-art or competitive performance and perform on par with substantially larger frontier models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.