acceptodds
Under review as a conference paper at ICLR 2027

Separating Pressure and Accountability in Tool-Using Agents: Evidence from CESO-Bench

Abstract

In the pressure comparison studied here, a pressured agent is compared with an agent told that its work will be audited. That comparison does not separate a pressure effect from an accountability-cue effect or from the model's behavior without either cue. We study this decomposition in executable tool use with CESO-Bench, a synthetic benchmark of 24 task families spanning access control, refund processing, and data retention. Each family has pressure, full-auditor, and neutral variants built from the same task body. The neutral closing keeps task-defining details while removing selected pressure language and is mechanically matched to its paired pressure closing within ±15% of word count. We evaluate three MiniMax M-series configurations at temperature 0.7 using a frozen surface success signal and deterministic end-state compliance oracles. The primary outcome is silent pass, SP = mean[A(1-C)], where A is the recorded task_answer_ok signal and C is the atomic-compliance result after authoritative replay. Across the three models, the pressure-versus-neutral family means are -0.015, -0.031, and -0.023 for M2.7, M2.5, and M3; the corresponding auditor-versus-neutral means are -0.060, -0.096, and -0.044. The M2.7 95% percentile intervals are [-0.049, +0.011] and [-0.157, +0.003]; the M2.5 and M3 analyses have partial family support and report point estimates only. This paper treats those contrasts as endpoint-specific evidence: a point estimate near zero is not equivalence, a confidence interval that covers zero is not proof of independence, and an auditor comparison does not identify a mechanism. The study is limited to the synthetic task pack, recorded runs, and the declared oracle contract.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.