acceptodds
Under review as a conference paper at ICLR 2027

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Abstract

Large language models often reason at length before answering, increasing cost and latency. We evaluate two reasoning interfaces in different model families under a shared design. Qwen3 receives concision and budget instructions, and gpt-oss uses its trained reasoning effort settings. These interfaces can shorten traces, but completed responses alone do not establish whether an interface improves earlier answers or changes when the model finishes. We evaluate paired runs at shared reasoning allowances and distinguish the answers models generate from the answers we read off their unfinished reasoning. The main comparisons cover Qwen3-14B and gpt-oss-20b/-120b on 198 GPQA Diamond and 500 MMLU-Pro questions. Across seven shared allowances, low gpt-oss effort improves generated-answer accuracy over high effort by 4.5-18.9 percentage points at 512 tokens, while high effort is 4.4-8.8 points more accurate at 8,192 tokens. The early advantage is concentrated in pairs where lower effort has completed and high effort remains active. For Qwen3-14B on MMLU-Pro, a concise instruction improves generated-answer accuracy by 3.5 points at 512 tokens and 2.5 points at 2,048. On pairs that both policies leave unfinished at 512 tokens, it still improves generated answers by 4.1 points. A budget instruction that announces the allowance shortens reasoning but gives smaller and less consistent gains, and matching the announced allowance to the evaluation horizon adds little in the tested comparisons. Neither the effort advantage nor the concise gain reflects adaptation to a stated allowance, and neither establishes uniformly better unfinished reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.