Learning What to Execute and When: Reasoning Protocols for LLM Inference
Abstract
Inference-time reasoning methods often extend a single language-model call through hand-designed procedures, such as voting, critique, program execution, or verification and revision. Learned multi-agent topologies optimize who communicates with whom, but have only matched hand-designed graphs under a common evaluation protocol. We introduce **REPROT**, which learns what to execute and when. REPROT represents a reasoning protocol as a stateful program over typed operations, specifying access to earlier results, operation composition, and runtime conditions that allocate extra computation to individual problems. Its language covers six common procedure families. A synthesizer examines search-set execution traces and proposes patches containing a diagnostic, a trigger condition, and a sub-protocol with defined visibility. A patch is accepted only if its net gain in problems fixed over problems broken passes a sign test on the full search set while meeting a mean-call budget. Protocol selection uses pooled search and disjoint validation data; the selected protocol is fixed before held-out testing. Across eight benchmarks, REPROT achieves the highest average accuracy, **84.2% at 3.9 calls per question**, compared with seventeen baselines reproduced under one protocol and the original ADAS and AFlow. The strongest hand-designed and learned-topology baselines score 79.7–80.3% at 5.2–11.0 calls; AFlow scores 82.6%. Paired comparisons with AFlow show four wins, four ties, and no losses. On TheoremQA and MuSR, REPROT scores 83.5% and 79.1%, versus 77.7% and 64.2% for the best reproduced baselines and 82.7% and 77.2% for AFlow. In both cases, the learned change is an instruction within a one- or three-call protocol. Free-code search finds a similar answer-format change on OlympiadBench. Condition overrides alter cost by up to 2.6×; forcing the first branch prevents LiveCodeBench escalation and lowers accuracy by 9.7 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.