Compare AI model performance across 48+ benchmarks. Our composite index aggregates coding, math, reasoning, and language scores into a single intelligence ranking.
At-a-glance rankings across the three dimensions that matter most.
Neura Intelligence Index; Higher is better
Max input tokens; Higher is better
USD per 1M output tokens; Lower is better
Ranked by the Neura Intelligence Index — a weighted composite of 48 benchmarks across 8 categories.
Find the sweet spot — models in the top-left quadrant offer the best value.
Top performers in each benchmark category.
Who leads on each individual benchmark — click any card to see full results.
Different tasks need different strengths. These indices re-weight our benchmarks for specific workflows.
Best for software development, code generation, and debugging
Best for scientific research, data analysis, and complex reasoning
Best for writing, editing, summarization, and creative tasks
Highest intelligence per dollar — the cost-efficiency sweet spot
The Neura Intelligence Index is a composite score (0-100) computed from 48+ individual benchmarks spanning 8 categories: coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks.
Not all models have scores on all benchmarks. The confidence indicator reflects benchmark coverage: high (>70% of benchmarks), medium (40-70%), or low (<40%). Weights are renormalized across available categories so models aren't penalized for missing data.
Scores are aggregated from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena. Each score includes a verification status (official, self-reported, or aggregated).
The Neura Intelligence Index is a composite score (0-100) that aggregates AI model performance across 15+ benchmarks spanning coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks. It uses min-max normalization and weighted category averaging to produce a single comparable score.
Benchmark scores are synced daily via an automated pipeline that aggregates data from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena.
Confidence reflects how many benchmarks a model has been tested on relative to the total. High means >70% benchmark coverage, medium is 40-70%, and low is <40%. Models with low coverage may have composite scores that shift as more benchmarks are added.
Category weights reflect real-world demand: Coding and Reasoning each get 20%, General and Math each get 15%, Language and Multimodal each get 10%, and Safety and Agent each get 5%. Weights are renormalized across available categories so models are not penalized for missing data.
| # | Model | Coverage | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
7— | Sora OpenAI🇺🇸1y | 89.4 | — | — | — | — | — | — | — | — | — | low |
8— | Veo 2 Google🇺🇸1y | 88.2 | — | — | — | — | — | — | — | — | — | low |
93— | Kling 1.6 Kuaishou🇨🇳1y | 64.2 | — | — | — | — | — | — | — | — | — | low |
127— | Runway Gen-3 Alpha Runway🇺🇸2y | 57.9 | — | — | — | — | — | — | — | — | — | low |
179— | Wan2.1 Alibaba🇨🇳1yOSS | 42.3 | — | — | — | — | — | — | — | — | — | low |
223— | Pika 2.0 Pika Labs🇺🇸1y | 32.3 | — | — | — | — | — | — | — | — | — | low |
246— | Minimax Video-01 MiniMax🇨🇳1y | 23.1 | — | — | — | — | — | — | — | — | — | low |
316— | LTX Video 13B Lightricks🇮🇱1yOSS | 3.8 | — | — | — | — | — | — | — | — | — | low |
Yes. Click any model to see its full benchmark profile, or use the comparison feature to compare up to 4 models side by side with radar charts and benchmark-by-benchmark scoring.