ELEMENTARY: SHERLOCKED by Causality — From Clues to Causes in Interactive LLM Agents
Abstract
Interactive LLM agents may succeed on the current task without recovering causal structure that remains useful when conditions or objectives change. We introduce ELEMENTARY, a generator-based escape-room benchmark that separates evidence sufficiency from agent reasoning failure by controlling how solution-relevant evidence becomes available within a shared interactive environment. The benchmark defines four evidence-conditioned reasoning regimes: Association exposes explicit relations; Observation provides sufficient passive evidence; Intervention requires active evidence acquisition; and Counterfactual requires reasoning from factual traces through historical revision and replay. Across twelve generator families, every generated question is validated so that the permitted evidence determines the required task-level solution. We evaluate seven proprietary and open-weight LLMs on 120 validated questions. No model is consistently strong across evidence conditions, and the highest average success rate is only %. More importantly, success on the current task can hide a failure to recover reusable causal structure. In I3 Causal Propagation, all model–question episodes reach the first control target, but only reaches the second target when successful control requires propagation-aware action ordering. Other diagnostics reveal additional failures in evidence search, passive inference, and historical revision. Together, these results show that current agents can often act on available clues without recovering the causal structure needed to succeed when the objective changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.