acceptodds
Under review as a conference paper at ICLR 2027

SeamBench: Evaluating Append-Only Model Handoffs in Code

Abstract

When an LLM service switches models mid-response, the replacement must finish code it did not write and cannot revise, and the client must decide whether to accept the result. We introduce SeamBench, a benchmark of 2,488 such append-only code handoffs from 428 HumanEval and MBPP tasks, built only from tasks both models solve on their own and labelled by executing the combined program. The acceptance decision turns out to be the hard part, and what a detector is allowed to see is a first-order design variable: in a controlled re-judging of 517 BigCodeBench handoffs, giving the judge the task description raises recall in every condition, while revealing where the switch occurred lowers AUC without the task and raises it with the task. Continuations fail on 10–26% of prefixes; 91% of failures are caught by syntax and entry-point checks, but the remaining semantic failures are caught by no rule and by a content-only judge fewer than half the time, at a false-rejection cost that varies from 2% to 42% across model directions. Switching itself is not the source of the failures: on 261 matched prefixes with fresh same-model restarts, replacing the continuation model never increases failure, and in one direction lowers it from 26% to 10%. All seams, labels and scorers are released under MIT and CC-BY licenses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.