Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...
| Benchmark | Score | Source |
|---|---|---|
| Chatbot Arena Elo | 1365 | |
| IFEval | 91.5 | |
| MMLU-Pro | 83.5 |
| Benchmark | Score | Source |
|---|---|---|
| HumanEval |
95.1 |
| SWE-bench Verified | 55.8 |
| Benchmark | Score | Source |
|---|---|---|
| BigBench-Hard | 93.2 | |
| GPQA Diamond | 74.8 |
| Benchmark | Score | Source |
|---|---|---|
| MMMU | 74.2 |
Anthropic