acceptodds
Under review as a conference paper at ICLR 2027

How Requirement Formats Affect Large Language Models for Code Generation: Empirical Study

Abstract

Large language models (LLMs) generate code from requirements expressed as prose, scenarios, tests, or specifications, yet the effect of these formats remains unclear. We compare 13 requirement formats across four programming benchmarks and five LLMs, then manually verify transformed requirements to examine whether rewriting errors explain the observed differences. Across model–benchmark pairs, the highest and lowest Pass@1 scores under automatic transformation differ by 6.34%–32.72%. Manual correction raises mean Pass@1 by 1.44%, but substantial differences among formats remain: semi-structured formats lead on average, while preferred formats vary by model and algorithmic task. We use these findings to develop ReForm, which estimates model-specific format preferences from execution outcomes grouped by algorithmic subcategory and selects a format for each new requirement. On held-out OJBench, ReForm achieves 30.00% mean Pass@1 across five models, compared with 22.84% for original requirements and 28.10% for the best single format measured on OJBench. For Claude-Sonnet-4.6, the selected formats also improve Pass@1 when combined with Self-Planning, SCoT, and OpenCode. These results show that requirement-format choice matters beyond transformation fidelity and that format selection transfers to new problems and downstream coding workflows. Our source code and details are publicly available at https://figshare.com/s/b8c47472426b87eb7776}{https://figshare.com/s/b8c47472426b87eb7776.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.