SeamBench: Evaluating Append-Only Model Handoffs in Code
Abstract
When an LLM service switches models mid-response, the replacement must finish code it did not write and cannot revise, and the client must decide whether to accept the result. We introduce SeamBench, a benchmark of 2,488 such append-only code handoffs from 428 HumanEval and MBPP tasks, built only from tasks both models solve on their own and labelled by executing the combined program. The acceptance decision turns out to be the hard part, and what a detector is allowed to see is a first-order design variable: in a controlled re-judging of 517 BigCodeBench handoffs, giving the judge the task description raises recall in every condition, while revealing where the switch occurred lowers AUC without the task and raises it with the task. Continuations fail on 10–26% of prefixes; 91% of failures are caught by syntax and entry-point checks, but the remaining semantic failures are caught by no rule and by a content-only judge fewer than half the time, at a false-rejection cost that varies from 2% to 42% across model directions. Switching itself is not the source of the failures: on 261 matched prefixes with fresh same-model restarts, replacing the continuation model never increases failure, and in one direction lowers it from 26% to 10%. All seams, labels and scorers are released under MIT and CC-BY licenses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.