acceptodds
Under review as a conference paper at ICLR 2027

AGC-Bench: Measuring Artificial General Creativity

Abstract

Creativity research has long debated whether creativity is a general ability or specific to particular domains (e.g., creative writing, visual art, scientific discovery), and whether it is separable from general intelligence. Both questions now apply to LLMs, but a fragmented evaluation landscape across hundreds of heterogeneous creativity benchmarks has left them empirically intractable. We introduce AGC-Bench, a meta-benchmark for artificial general creativity built from a PRISMA-compliant systematic review of the AI creativity literature ( papers screened, unique benchmarks identified) paired with an agentic onboarding harness that converts source-paper benchmarks into runnable HELM-style scenarios. The first release covers datasets ( text-only, multimodal) spanning brainstorming, problem solving, STEM, narrative, figurative language, and humor. To address bias in LLM-as-judge, we apply Judge Response Theory (JRT)—a psychometric calibration of judge leniency/severity—with three frontier LLM judges. We then fine-tune Qwen3-30B-A3B-Instruct-2507 on the resulting JRT-corrected ratings to produce AGC-Judge, an open-weight scoring model that matches the three-judge ensemble and predicts frontier-judge ratings with high accuracy on creativity benchmarks it was not trained on. Results reveal frontier models at the top of the leaderboard, with open-weight models close behind. However, different models show different creative strengths, ranking higher on some domains (e.g., creative writing) than others (e.g., scientific ideation). We conduct several experiments and report three main findings. First, applying factor analysis across LLMs, we recover a single creativity factor ‘c’, analogous to the ‘g’ factor of general intelligence, that explains % of variance, related to but separable from general knowledge and reasoning. Second, in within-model comparisons, we show that prompting models to “be creative” boosts their performance far more than enabling their reasoning capability, evidence that the benchmark tracks creativity specifically. Third, on a human-matched subset of five tasks, the top human still leads the top LLM on the novelty and diversity composite. We release AGC-Bench with a public leaderboard, the onboarding harness, AGC-Judge weights, and human comparison data as open infrastructure for measuring AI creativity at scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.