CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection
Abstract
Evaluating Large Language Models (LLMs) incurs substantial computational overhead during continuous development, since a fixed benchmark is rerun for many checkpoints. Coreset selection reduces this workload, but existing methods either require historical response logs that new benchmarks lack, as in Item Response Theory, or inherit the lexical bias of question-only embeddings. We propose CoT-Core, a training-free core question selection framework. Because lexically disparate questions can require related reasoning operations, CoT-Core prompts an LLM to unroll a zero-shot Chain-of-Thought (CoT) trajectory for each question and clusters the resulting question-trajectory representations to choose representative items. The subset is constructed once, without reference answers or historical target-model responses, and reused across evaluations. On GSM8K, MMLU, MMLU-Pro, and GPQA, CoT-Core attains the strongest history-free score estimation in most dataset-budget settings. Rerunning the full evaluation pipeline for eleven target models on MMLU-Pro further shows that a 5% subset reduces output tokens by 94.77%, while its estimated scores deviate from full-benchmark scores by only 0.0065 on average. CoT-Core therefore provides a practical option for repeated evaluation when historical responses are unavailable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.