Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 99.8 | verified |
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 91.3 | verified |
| ARC-AGI v2 | 68.8 | verified |
| Humanity's Last Exam | 53.1 | verified |
| Benchmark | Score | Source |
|---|---|---|
| CharXiv Reasoning | 77.4 | verified |
| MMMU-Pro | 77.3 | verified |
| Benchmark | Score | Source |
|---|---|---|
| BrowseComp | 84 | verified |
| OSWorld | 72.7 | verified |
| MCP Atlas | 62.7 | verified |
Anthropic