Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
| GPQA Diamond | 70.4 | verified |
| Benchmark | Score | Source |
|---|---|---|
| ScreenSpot Pro | 60.5 | verified |
| MMMU-Pro | 60.4 | verified |
| CharXiv Reasoning | 48.9 | verified |
| Benchmark | Score | Source |
|---|---|---|
| OSWorld | 30.3 | verified |
Qwen