acceptodds
Under review as a conference paper at ICLR 2027

AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

Abstract

LLM agents are increasingly taking actions with real consequences, yet safety evaluations often reduce their behavior to a final safe or unsafe outcome, obscuring whether the agent recognized the risk and how it responded. We introduce AURA-Eval, a framework for evaluating whether agents recognize contextual risk and adapt their behavior accordingly. AURA-Eval truncates tool-use trajectories at safety-critical decision points and asks the evaluated model to generate the next thought and action. It creates controlled scenario variations along six risk mechanisms: harm intensity, scale, target susceptibility, system dependency, oversight, and reversibility. It also creates paired scenarios that preserve the underlying goal but differ in whether safe fulfillment is available, allowing unsafe compliance and over-refusal to be evaluated together. An LLM-judge panel, validated against human annotations, evaluates each continuation along three axes: risk detection from the reasoning trace, and action type and safety from the action. From 157 R-Judge seed trajectories, we construct 1,249 evaluation items and evaluate 20 frontier, open-weight, and safety-tuned models. Across all 20 models, unsafe-action rates are higher when no safe path for request fulfillment is available than when one is; for frontier models, the rates are 27-40% versus 5-7%. When a safe path exists, frontier models rarely refuse; when none exists, they most often propose an alternative. Open-weight models, by contrast, more often execute the request than propose alternatives and act unsafely on 58-83% of items. Controlled risk-mechanism analyses further show particularly high unsafe rates under changes to oversight and scale; removing pre-execution oversight flips 52.3% of cases that were safe on the corresponding original scenario to unsafe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.