acceptodds
Under review as a conference paper at ICLR 2027

Verifier-Native Control and Trajectory Evaluation for Operations Research Repair Agents

Abstract

Operations Research (OR) repair is iterative: a candidate edit is solved, checked, diagnosed, and revised. We study how LLM agents use this feedback while preserving operational validity. **Repair-Sufficient State (RSS)** compiles solver and domain-checker observations into typed failure state; **Counterfactual Verifier Loop (CVL)** organizes probes and repairs around verifier transitions; and **OR-INSPECT-BENCH** evaluates fixed-trace diagnosis and live repair. Across 19 directly measured model snapshots, the full pipeline improves strict recovery over one-shot prompting. Matched controls show that RSS improves on raw multi-turn feedback across three anchors. In a fresh same-window Qwen comparison, adding reference-free output normalization to RSS raises strict all-error recovery from 30.6% to 42.8%; the full CVL-r2 configuration reaches 45.8%, a further 3.1-point estimate. A Qwen factorial also finds that canonical-emission correction raises strict recovery by 14.4 and 20.6 points with probes enabled and disabled. Across six paired model anchors, additional control gains vary by model. OR-INSPECT-BENCH combines fixed-trace diagnosis, live repair, and composition holdouts. Together, the results quantify contributions from typed verifier state and output handling, and characterize when added control improves recovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.