CodeInFlight: Can LLMs Code Well under Changing User Intentions?
Abstract
LLM-based coding agents increasingly work alongside users who revise their requests before implementation is complete. Such revisions can invalidate an agent’s current plan or edits, requiring it to redirect unfinished work while preserving progress that remains useful. However, existing repository-level benchmarks largely assume fixed specifications or introduce new requirements only between completed turns, and therefore do not directly evaluate adaptation to changes that arrive during execution. To address this gap, we introduce CodeInFlight, a dataset and repository-level benchmark for evaluating how coding agents adapt to changing user requests while they work. We curate 1,551 execution-validated source tasks spanning 534 repositories and ten programming languages, from which we derive a stratified evaluation set of 287 interaction instances covering requirement changes, cancellation with replacement, and additional tasks. For each instance, UPFRONT, SEQUENTIAL, and IN-FLIGHT reveal the final request before coding, after the initial attempt, or during execution, respectively. Comparing each of 19 configurations with itself, we find that all six configurations resolving at least half of the instances when a change follows their first attempt resolve fewer when it arrives during execution. They generate 11% fewer output tokens but resolve 3.9 points fewer instances, for both withdrawals and additions. In paired runs of eight open-weight models, thinking raises resolution under both deliveries by similar amounts. CodeInFlight thus indicates that these higher-performing configurations adapt less reliably to changes during execution than to changes after completion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.