ARC: Learning Selective Answer Repair from Matched Outcomes
Abstract
Small language models offer efficient inference, but errors in their initial answers can limit their reliability on reasoning tasks. Self-refinement and external tools can help correct these errors, although applying repairs to every answer adds computation and can damage correct answers. The challenge is to predict which repair, if any, will improve an answer before executing it. We introduce ARC (Anchor-conditioned Repair Controller), which learns this decision from matched repair outcomes. During training, we apply all available repairs to the same initial output, called the Anchor, and use their final-answer scores to supervise retention and repair selection. At deployment, an output-conditioned hierarchical router reads the question, generated answer, and public reasoning, then retains the Anchor or executes one repair with the base model fixed. On Qwen3-8B across five reasoning task families, the main evaluation yields a mean task score of 0.6915, a 6.75% improvement over retaining initial answers. It uses 0.155 additional LLM calls and 0.348 s of recorded LLM request time per example, reducing these costs by 84.5% and 90.9%, respectively, compared with one self-refinement call per answer. A separate deterministic evaluation over three seeds yields nearly equal mean in-distribution scores with 24.8% fewer additional calls than routing without the generated answer or reasoning. Experiments with the other evaluated base models also show improvements over answer retention and question-only routing. Our code is available at https://anonymous.4open.science/r/ARC-474B/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.