Preprint
Large Language Models

Has GPT-5 Achieved Spatial Intelligence?

Zhongang Cai, Yubo Wang, Qingping Sun, Ruisi Wang, Chenyang Gu, Wanqi Yin, Zhiqian Lin, Zhitao Yang, Chen Wei, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Jiaqi Li, Xiangyu Fan, Hanming Deng, Lewei Lu, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, Lei Yang
August 18, 202514 citations

14

Citations

2

Influential Citations

Venue

2025

Year

Abstract

Multimodal models have achieved remarkable progress in recent years. Nevertheless, they continue to exhibit notable limitations in spatial understanding and reasoning, the very capability that anchors artificial general intelligence in the physical world. With the recent release of GPT-5, allegedly the most powerful AI model to date, it is timely to examine where the leading models (GPT, Gemini, Grok, Seed, Qwen, and Intern) stand on the path toward spatial intelligence (SI). We thus propose EASI for holistic Evaluation of multimodAl LLMs on Spatial Intelligence. EASI conceptualizes a comprehensive taxonomy of spatial tasks that unifies existing benchmarks and a growing collection of newly curated ones, enabling systematic evaluation of state-of-the-art models. In this report, we conduct the study across eight key benchmarks, at a cost exceeding ten billion total tokens. Our empirical study then reveals that (1) GPT-5 demonstrates unprecedented strength in SI, yet (2) still falls short of human performance significantly across a broad spectrum of SI-tasks. Moreover, we (3) show that SI-tasks expose greater model capability deficiency than non-SI tasks, to the extent that (4) proprietary models do not exhibit a decisive advantage when facing the most difficult ones. In addition, we conduct a qualitative evaluation across a diverse set of scenarios that are intuitive for humans, yet fail the most advanced multimodal models. EASI is an ongoing community effort: we have open-sourced the EASI codebase that provides a one-stop and reproducible solution with standardized interfaces, integrated protocols and prompts that significantly reduce the friction of configuring and running multiple benchmarks; we have also launched an accompanying EASI leaderboard to provide a continually updated snapshot of model performance across the full SI spectrum, accelerating collective progress toward robust SI.

Analysis

Why This Paper Matters

Spatial intelligence is a cornerstone of artificial general intelligence, enabling agents to understand and interact with the physical world. Despite rapid advances in multimodal LLMs, systematic evaluation of spatial reasoning has been fragmented. This paper addresses that gap by introducing EASI, a holistic benchmark that unifies existing spatial tasks and adds new ones. The timing is critical: with GPT-5's release, the community needs an objective measure of progress. The study's scale—over ten billion tokens—provides robust evidence that even the most advanced models still lack human-level spatial understanding.

Technical Contributions

  • EASI Taxonomy: A comprehensive classification of spatial tasks that integrates prior benchmarks and newly curated scenarios, enabling systematic comparison.
  • Large-Scale Evaluation: Tests six leading models (GPT, Gemini, Grok, Seed, Qwen, Intern) across eight benchmarks with massive token consumption, ensuring statistical reliability.
  • Open-Source Infrastructure: Released codebase with standardized interfaces, protocols, and prompts, plus a live leaderboard for ongoing community tracking.
  • Qualitative Analysis: Identifies intuitive human scenarios where all advanced models fail, highlighting specific weaknesses.

Results

  • GPT-5 demonstrates the strongest spatial intelligence among all tested models, but still significantly underperforms humans across a broad spectrum of tasks.
  • Spatial tasks expose larger capability deficiencies than non-spatial tasks, suggesting spatial reasoning is a key bottleneck.
  • Proprietary models do not show a decisive advantage on the most difficult spatial tasks, indicating room for open-source models.
  • The qualitative evaluation reveals failure modes in scenarios that are trivial for humans, such as understanding object relationships in 3D space.

Significance

This work sets a new standard for evaluating spatial intelligence in AI. By providing a unified benchmark and leaderboard, it enables the community to track progress systematically. The finding that even GPT-5 lags behind humans underscores the distance to AGI. The open-source tools lower barriers for researchers, accelerating collective progress. As spatial reasoning is critical for robotics, autonomous driving, and AR/VR, this benchmark will likely influence future model development and deployment.