Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 68 | verified |
| TerminalBench | 40.5 | verified |
| Benchmark | Score | Source |
|---|---|---|
| AIME 2025 | 93.9 | verified |
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 81 | verified |
| Humanity's Last Exam | 17.2 | verified |
| Benchmark | Score | Source |
|---|---|---|
| BrowseComp | 45.1 | verified |
Z.ai