acceptodds
Under review as a conference paper at ICLR 2027

TERSE: Trace-Efficiency Reward via Squeezing rEdundancy for Efficient Large Reasoning Models

Abstract

Large reasoning models perform well on complex problems by generating long chains of thought, but these traces are often redundant, raising inference cost and latency without improving accuracy. Reinforcement learning methods curb this by penalizing length, which captures how much a model writes but not how much of it merely repeats: two correct traces of equal length can carry very different amounts of redundancy. Prior methods address this by adding a second signal, derived from external information such as a reflection lexicon, an auxiliary model, or the model's internals. We instead read a second signal from the trace itself, the gzip compressibility of its raw bytes, which depends on bytes rather than meaning and is therefore not tied to a domain or model family. Building on it, we propose Trace-Efficiency Reward via Squeezing rEdundancy (TERSE), a reward for group-relative reinforcement learning. TERSE residualizes compressibility against length so that it rewards only the redundancy length does not already explain, and standardizes both signals within difficulty buckets to remain stable on hard prompts where correct rollouts are scarce. On five mathematical reasoning benchmarks with DeepSeek-R1-Distill-Qwen at 1.5B and 7B, improves accuracy over the base model by 7.8 and 3.4 points while cutting response length by 54.4% and 37.4%, improving the accuracy–length trade-off at both scales.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.