acceptodds
Under review as a conference paper at ICLR 2027

CodeInFlight: Can LLMs Code Well under Changing User Intentions?

Abstract

LLM-based coding agents increasingly work alongside users who revise their requests before implementation is complete. Such revisions can invalidate an agent’s current plan or edits, requiring it to redirect unfinished work while preserving progress that remains useful. However, existing repository-level benchmarks largely assume fixed specifications or introduce new requirements only between completed turns, and therefore do not directly evaluate adaptation to changes that arrive during execution. To address this gap, we introduce CodeInFlight, a dataset and repository-level benchmark for evaluating how coding agents adapt to changing user requests while they work. We curate 1,551 execution-validated source tasks spanning 534 repositories and ten programming languages, from which we derive a stratified evaluation set of 287 interaction instances covering requirement changes, cancellation with replacement, and additional tasks. For each instance, UPFRONT, SEQUENTIAL, and IN-FLIGHT reveal the final request before coding, after the initial attempt, or during execution, respectively. Comparing each of 19 configurations with itself, we find that all six configurations resolving at least half of the instances when a change follows their first attempt resolve fewer when it arrives during execution. They generate 11% fewer output tokens but resolve 3.9 points fewer instances, for both withdrawals and additions. In paired runs of eight open-weight models, thinking raises resolution under both deliveries by similar amounts. CodeInFlight thus indicates that these higher-performing configurations adapt less reliably to changes during execution than to changes after completion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.