acceptodds
Under review as a conference paper at ICLR 2027

Agent Carcinization: A Testable Framework for Architectural Convergence and Control Trade-offs in LLM Agents

Abstract

LLM agents increasingly rely on external runtime controls to constrain actions, track state, and support recovery. We use agent carcinization to frame the hypothesis that shared operational pressures favor recurring control structures, and ask how their benefits should be evaluated. We study evaluation contracts: the rules that determine which task outcomes, procedural requirements, and boundary constraints count as success. Re-scoring a fixed corpus of 200 episodes across four models, four tasks, and five runtime conditions reveals substantial endpoint sensitivity. On three business tasks, the model-balanced difference between a gated runtime and a placebo control changes from +27.08 percentage points under the original evaluator to −29.86 under an outcome-only ledger, −13.19 when completion is additionally required, and −2.08 when recorded boundary violations are also excluded. These comparisons hold execution trajectories fixed, isolating the effect of scoring rules rather than changes in agent behavior. External comparisons provide additional task context but do not reproduce the reversal, while targeted cancellation experiments demonstrate blocking and recovery under specified conditions. We contribute executable counterexamples, decomposed endpoints, and replayable evidence for auditing evaluation contracts. Our findings show why claims about agent runtime structure should distinguish task outcomes, procedural compliance, and boundary enforcement, while making endpoint definitions and their evidential limits explicit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.