acceptodds
Under review as a conference paper at ICLR 2027

TRUST: Testing Reliability Under Situational Tensions in Coding Agents

Abstract

Coding agents have advanced rapidly alongside large language models and can now autonomously handle increasingly complex software engineering tasks, marking an important step toward artificial general intelligence (AGI). As these agents become more capable, reliable evaluation of their strengths and limitations has become an essential challenge. Existing coding agent benchmarks largely measure performance through aggregate success rates on common software engineering tasks, offering valuable assessments of coding capability. Yet aggregate success can miss rare but consequential failures in agent behavior. An agent may perform strongly on standard benchmarks while still violating user intent, crossing engineering boundaries, or making decisions that are not supported by sufficient evidence in edge scenarios. In real-world use, such failures can directly affect user trust and practical reliability. We introduce **TRUST**, a benchmark for **T**esting **R**eliability **U**nder **S**ituational **T**ensions in Coding Agents, to evaluate this often-overlooked aspect of agent behavior. TRUST combines evolved public coding benchmarks with expert-designed scenarios grounded in real software engineering practices. Each instance uses process-oriented rubrics to identify specific behavioral failure modes in the execution trajectory, independent of the final task outcome. TRUST contains 20 targeted stress tests across six dimensions of behavioral reliability. Our evaluation protocol specifies three runs per agent–task pair to examine run-to-run behavioral variability and help reveal intermittent failures. By exposing reliability risks hidden by aggregate task success, TRUST complements capability-focused evaluation with a direct assessment of agent behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.