acceptodds
Under review as a conference paper at ICLR 2027

Reflective Gradients: Feasibility–Task Tradeoffs in Solver Distillation with Verified Action Release

Abstract

Neural policies can violate explicit action constraints, while repairing every proposal with an optimization solver adds inference cost. We study whether detached solver supervision reduces repair dependence when an independent verifier remains responsible for action release. Admissible actions are finite unions of bounded polyhedra. The verifier checks proposals, requests repair when needed, and rechecks the decoded serialized action under the current state and rules. Under explicit compilation, arithmetic, state consistency, and consumer assumptions, every released action satisfies the encoded constraints independently of solver status. Feasibility verification alone gives no general approximation guarantee for branch-candidate selection; we reproduce severely suboptimal targets from a numerical repair service. A locally frozen study uses 20 paired trials per setting across disconnected and convex numerical tasks and a synthetic resource allocation application. Exact teachers, task reweighting, branch models, and fresh direct task optimization expose distinct feasibility, utility, and computational tradeoffs. Replacing numerical corrections with exact targets changes those tradeoffs without demonstrating a consistent learning advantage over task reweighting. Regret curves reveal threshold-dependent rankings, and the allocation task exposes poor learned utility at the fixed training budget. We provide conditional correctness proofs, analytic counterexamples, and auditable computation records. The findings motivate specifying teacher quality and evaluating complete decision cost under explicit task requirements rather than treating fewer repairs as sufficient evidence of useful constraint internalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.