acceptodds
Under review as a conference paper at ICLR 2027

Who Is at Risk? Reading LLM Decision Biases into Evacuation Simulations

Abstract

In a tsunami, who survives depends on when and how each resident decides to leave. Language models now offer a way to supply these decision traits, such as how strongly a resident relies on an official warning, to every resident of an evacuation simulation, and such simulations are usually judged by whether the death toll looks plausible. This test cannot tell whether each resident received the right traits: when residents are placed at random, shuffling their traits leaves the expected death toll unchanged but changes who dies. We therefore propose checks on who holds which trait: whether shuffling changes who dies, whether the groups that fare worst match evidence from real evacuations, whether a direction read from the model does better than random directions, and whether different models agree. The checks need no data on individual behavior, and they catch errors we inject on purpose. We read six decision biases from the hidden states of four open models into two established tsunami evacuation simulators. Shuffling the scores barely changes the death toll but changes about half of the victims, and sources with similar death tolls differ on the checks. Warning reliance is consistent across models while herd influence is not, and all four models rate older residents as the most warning-reliant, the opposite of 2011 mortality. A plausible death toll, and even agreement between models, can hide errors in who is put at risk.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.