acceptodds
Under review as a conference paper at ICLR 2027

GENIUS: Generative Fluid Intelligence Evaluation Suite

Abstract

Unified Multimodal Models (UMMs) have shown remarkable progress in visual generation. Yet, existing benchmarks predominantly assess *Crystallized Intelligence*, which relies on recalling accumulated knowledge and learned schemas. This focus overlooks *Generative Fluid Intelligence (GFI)*: the capacity to induce patterns, reason through constraints, and adapt to novel scenarios on the fly. To rigorously assess this capability, we introduce **GENIUS** (**GEN**erative Fluid **I**ntelligence Eval**U**ation **S**uite). We operationalize *GFI* through three complementary dimensions. These include *Inducing Implicit Patterns* (e.g., inferring personalized visual preferences), *Executing Ad-hoc Constraints* (e.g., visualizing abstract metaphors), and *Adapting to Contextual Knowledge* (e.g., simulating counter-intuitive physics). Collectively, these dimensions challenge models to solve problems whose governing rules are established by the immediate context. Our systematic evaluation of 14 representative models reveals significant performance deficits in these tasks. Our diagnostic analysis reveals a gap between recognizing the intended outcome and faithfully generating it: models can identify the expected output in a discriminative VQA probe, yet struggle to execute contextual rules during generation. To bridge this gap, we propose a training-free attention intervention strategy. Ultimately, **GENIUS** establishes a rigorous standard for *GFI*, guiding the field beyond knowledge utilization toward dynamic, general-purpose reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.