The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...
| Benchmark | Score | Source |
|---|---|---|
| FrontierMath | 5.5 | verified |
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 78 | verified |
| Benchmark | Score | Source |
|---|---|---|
| MMMU | 77.6 | verified |
| Benchmark | Score | Source |
|---|---|---|
| TAU-Bench Retail | 70.8 | verified |
OpenAI