acceptodds
Under review as a conference paper at ICLR 2027

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Abstract

Evaluating Large Language Models (LLMs) incurs substantial computational overhead during continuous development, since a fixed benchmark is rerun for many checkpoints. Coreset selection reduces this workload, but existing methods either require historical response logs that new benchmarks lack, as in Item Response Theory, or inherit the lexical bias of question-only embeddings. We propose CoT-Core, a training-free core question selection framework. Because lexically disparate questions can require related reasoning operations, CoT-Core prompts an LLM to unroll a zero-shot Chain-of-Thought (CoT) trajectory for each question and clusters the resulting question-trajectory representations to choose representative items. The subset is constructed once, without reference answers or historical target-model responses, and reused across evaluations. On GSM8K, MMLU, MMLU-Pro, and GPQA, CoT-Core attains the strongest history-free score estimation in most dataset-budget settings. Rerunning the full evaluation pipeline for eleven target models on MMLU-Pro further shows that a 5% subset reduces output tokens by 94.77%, while its estimated scores deviate from full-benchmark scores by only 0.0065 on average. CoT-Core therefore provides a practical option for repeated evaluation when historical responses are unavailable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.