Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
| Benchmark | Score | Source |
|---|---|---|
| MMMLU | 89.1 | verified |
| Benchmark | Score | Source |
|---|---|---|
| TerminalBench | 50 | verified |
| Benchmark | Score | Source |
|---|
| AIME 2025 | 87 | verified |
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 83.4 | verified |
| Benchmark | Score | Source |
|---|---|---|
| TAU-Bench Retail | 86.2 | verified |
| OSWorld | 61.4 | verified |
Anthropic