GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
| Benchmark | Score | Source |
|---|---|---|
| MMMLU | 87.3 | verified |
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 54.6 | verified |
| Benchmark | Score | Source |
|---|
| AIME 2025 | 46.4 | verified |
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 66.3 | verified |
| Humanity's Last Exam | 5.4 | verified |
| Benchmark | Score | Source |
|---|---|---|
| MMMU | 74.8 | verified |
| CharXiv Reasoning | 56.7 | verified |
| Benchmark | Score | Source |
|---|---|---|
| TAU-Bench Retail | 68 | verified |
OpenAI