acceptodds
Under review as a conference paper at ICLR 2027

Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs

Abstract

Accuracy per token summarizes reasoning efficiency but does not reveal whether differences arise from failed attempts, costly successful answers, or disproportionately long reasoning relative to the effort a problem demands. We introduce a two-layer protocol that connects these distinctions to the same efficiency score. The outcome layer separates answer availability, correctness among available answers, and generation cost. The workload layer relates cost to task structure, separating mean generated tokens per workload unit from their association with workload. Across 785,541 responses on twelve benchmarks, cost gaps persist among successful answers: Qwen3.5-27B generates 1.79 to 3.46 times as many tokens as Gemma-4-31B for correct outputs on shared problems across ten benchmarks. Relative costs also change with workload. On ZebraLogic, Qwen3.5-27B's cost relative to GPT-OSS-120B changes from 4.49 at low workload to 0.78 at high workload, despite similarly long responses on average. Across nine checkpoints on four annotated benchmarks, simultaneous 95% confidence intervals support relative-cost changes in 82 of 144 model-pair comparisons and effort-induced changes in 21 of 36 within-checkpoint comparisons. These results show that successful-answer cost and workload response are complementary properties of reasoning models: cost gaps can persist on shared successful questions while changing substantially with task structure and inference settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.