Benchmarking Embodied Fluid Intelligence
Abstract
We study fluid intelligence, the capacity to efficiently infer and exploit the actionable structure of an unfamiliar environment from limited action–perception experience, in interactive 3D worlds. We present Flint, a benchmark of 180 tasks, validated for solvability, unfamiliarity, and meaningful controlled variation. Each task is an unfamiliar puzzle in a physics-enabled 3D room. Instead of manually creating this benchmark, we design a framework that uses coding agents to automatically create such tasks. In particular, we ask coding agents to author task generators that allow us to generate not only individual tasks but their controlled variants. Evaluating frontier multimodal LLM agents on Flint, we find that GPT-6 Astra and Claude Opus 5.5 have increased the solve rate by > 35 percentage points over prior models. Yet, while they saturate many other benchmarks, they leave more than a third of the tasks unsolved on Flint. Our analysis sheds light on the challenges and opportunities for advancing embodied fluid intelligence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.