ClarifyCodeBench: Evaluating LLMs on Clarifying Underspecified Requirements for Code Generation
Abstract
Large language models (LLMs) have emerged as powerful programming assistants. However, the efficacy of code generation is fundamentally constrained by the quality of input requirements, which, in real-world software development, are frequently underspecified: they omit information needed to determine the intended behavior. Although LLMs excel at one-shot code synthesis, their ability to proactively clarify user intent remains underexplored despite its importance to robust software engineering. Existing benchmarks largely overlook this interactive bottleneck, using perfectly specified prompts that do not reflect the iterative nature of requirement elicitation. To bridge this gap, we introduce ClarifyCodeBench, a novel interactive benchmark specifically designed to evaluate LLMs' capability in resolving requirement underspecification. Constructed from real world programming tasks, ClarifyCodeBench features high-quality manual annotations, including 11 unique underspecification types, associated clarification questions, and corresponding ground-truth answers. Furthermore, we formalize two rigorous metrics to assess the interaction quality: Key Question Rate (KQR), which measures the fraction of key questions that a model asks about, and Turn-discounted Key Question Rate (TKQR), which also penalizes inefficient questioning. We conduct a systematic evaluation of four current LLMs using ClarifyCodeBench. Our empirical results yield three critical insights: 1) Capability Decoupling: Strong code generation performance does not inherently translate to strong requirement clarification efficacy; 2) The Reasoning Paradox: While increasing computation in reasoning models enhances code correctness, it reduces clarification coverage; 3) The Multi-underspecification Ceiling: LLMs' clarification performance degrades sharply as the density of underspecification increases, revealing a significant bottleneck in handling complex, real-world specifications. Our work emphasizes the necessity for future research to move from static synthesis to interactive requirement elicitation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.