acceptodds
Under review as a conference paper at ICLR 2027

THINKING FAST WHEN YOU HAVE NO TIME: TRAINING REASONING LLMs UNDER TIME CONSTRAINTS

Abstract

Reasoning language models achieve strong accuracy by generating long chain- of-thought (CoT) traces, but these traces dominate inference cost. Prior work reduces this cost using explicit mechanisms such as length penalties, early-exit mechanisms, or distilled shorter traces that constrain output length. In this work, we hypothesize that shorter responses can emerge implicitly from the structure of the training environment by grouping multiple questions into a single prompt together under a shared response budget (measured in response tokens) without any explicit length-based objective. We test this hypothesis by post-training lan- guage models using reinforcement learning to solve multi-question quizzes. We construct GSM8K-Quiz, Math-Quiz, and SVAMP-Quiz by grouping questions from established datasets into five-question prompts. Crucially, the budget is not uniformly enforced for each question, so the policy is free to spend more reasoning on harder questions and less on easier ones without any explicit difficulty super- vision. We train Qwen2.5-Math-1.5B-Instruct with proximal policy optimization (PPO), separately on GSM8K-Quiz and Math-Quiz. In quiz evaluation, under shared response budgets, the quiz-trained PPO policies use about 50% fewer tokens than the base instruct model on GSM8K-Quiz and 35% fewer on Math-Quiz, with comparable accuracy. In standard single-question evaluation, the efficiency gains persist: on the GSM8K test set, quiz-trained policies use 9–10% fewer tokens than PPO policies trained on single questions, again with comparable accuracy. Our results demonstrate that training in an environment with implicit time pressure can be sufficient to induce compute-efficient reasoning strategies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.