acceptodds
Under review as a conference paper at ICLR 2027

AmbiguityBench: Reasoning Modes Bidirectionally Shift Ambiguity Sensitivity in Frontier LLMs

Abstract

We introduce AmbiguityBench, a diagnostic benchmark for evaluating how large language models behave under ambiguity (unknown probabilities) versus risk (known probabilities), and use it to audit reasoning-enabled configurations across providers. Four probe families—Ellsberg variants, matched Allais risk controls, and tool-selection scenarios—are administered under direct, chain-of-thought, and ambiguity-aware prompting to four base/reasoning pairs from Google, DeepSeek, and Anthropic, plus an open-weight check. Across 23,560 trials, unprompted (direct) elicitation yields large, provider-dependent shifts: Claude Sonnet 4.6 thinking reduces the Ambiguity Premium from +0.396 to +0.080 (delta = +0.316, q_FDR = 0.002; +0.344 at matched temperature, p < 0.0001), while DeepSeek-R1 increases it (delta = -0.111, q = 0.002). Both Gemini generations are small or null. These pair differences largely disappear under chain-of-thought: CoT on Claude base already matches thinking-mode AP, and all four CoT pair tests are non-significant after FDR (q > 0.68). Matched Allais controls remain flat (8/8 cells q > 0.29). Ambiguity sensitivity is therefore a real, model-specific behavioral property, but bidirectionality is a fact about unprompted choice, not a general law of reasoning APIs. We release probes, data, and code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.