Single-Configuration Evaluation Can Mislead Learned MILP Cut Selection: Seed Fragility, Measurement-Era Artifacts, and the Limits of Score Blending
Abstract
Learned components inside mixed-integer linear programming (MILP) solvers are often evaluated under a single experimental configuration, leaving the reliability of reported gains under solver and system variability unclear. We study this problem in the setting of learned cut selection in SCIP. Re-evaluating a frozen supervised cut scorer across six permutation seeds, we find substantial but sign-unstable effects on the search trajectory: only 2 of 31 test instances with complete six-seed data retain one benefit sign across all seeds, and pairwise sign agreement is 0.52. Two additional LightGBM cut scorers, trained with imitation and multiple-instance-learning targets, exhibit nearly identical sign instability (pairwise agreement 0.49 and 0.48, with no instance unanimous across all six seeds), ruling out our attribution-label recipe as the sole cause. Under a symmetric noise model, this disagreement places an approximately 60% upper bound on the fresh-seed sign accuracy of any seed-agnostic instance gate; the gates we test fail consistently with this ceiling. We uncover a second, orthogonal source of unreliability: an earlier 31/0 win–loss result against default SCIP was an artifact of a baseline measured under machine load. CPU-time accounting alone does not remove this confound, because memory pressure can inflate CPU cost per solver iteration; load logging and per-iteration rate probes expose it. After clean re-measurement, score blending is consistently less fragile than replacing SCIP's native score, but it does not produce robust downstream improvement: on benchmark-suite instances the intervention perturbs trajectories while about half of runs reach the same final primal and dual bounds as default, and on harder held-out instances the remaining positive structure is concentrated at intermediate difficulty and remains seed-dependent. These results motivate a two-axis evaluation protocol for learned solver components: vary solver symmetry-breaking conditions and verify the machine state under which paired comparisons are measured.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.