Enterprise’s Last Exam: A Benchmark for Organizational Judgment under Fragmented Enterprise Context
Abstract
Frontier language models perform strongly on benchmarks of coding, mathematics, knowledge, and general reasoning. However, enterprise deployments often require a different capability: making decisions whose correctness depends on organization-specific entities, historical precedent, conflicting systems of record, changing policy versions, distributed or informal authority, and temporal context. We introduce Enterprise's Last Exam (ELE), a benchmark for this capability, which we term organizational judgment. ELE complements benchmarks of generic knowledge, artifact production, and software interaction by directly measuring whether models identify and correctly apply the organization-specific evidence that governs a decision. ELE comprises 630 practitioner-authored scenarios across three complementary model-blind evaluation sets: a Core set (n=384), a difficulty-focused Challenge set (n=156), and 90 counterfactual scenarios forming 45 paired cases in which a single decision-critical fact changes. Scenarios span six dimensions: entity resolution under ambiguity, precedent-based exception handling, cross-system synthesis, policy-version reasoning, approval-chain reconstruction, and temporal decision consistency. We evaluate seven proprietary and open-weight models under a controlled direct-prompt condition at temperature zero. On the primary Core set, accuracy ranges from 77.4% to 97.4%; no model saturates the benchmark. Performance varies sharply by reasoning type: models are strongest on entity resolution, approval-chain reconstruction, and temporal consistency, but substantially weaker when they must apply precedent or reconcile conflicting evidence across systems; cross-system synthesis ranges from 67% to 90%. Challenge accuracy ranges from 73.7% to 87.8%. On the 45 counterfactual pairs, the fraction for which models answer both variants correctly ranges from 62% to 89%, indicating substantial variation in robustness when a single governing organizational fact changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.