Qwen2-0.5B|1.5B|7B|57B-A14B-MoE|72B logo

Qwen2-0.5B|1.5B|7B|57B-A14B-MoE|72B

Free

Open-source LLMs with 128K context, multilingual support, and state-of-the-art performance

FreeFree tier
Inputs: textOutputs: text
Type
Open Source
Company
Alibaba Cloud (Qwen Team)

About Qwen2-0.5B|1.5B|7B|57B-A14B-MoE|72B

Qwen2 is the evolution from Qwen1.5, a series of open-source language models developed by the Qwen Team (Alibaba Cloud). It includes pretrained and instruction-tuned models in five sizes: Qwen2-0.5B, Qwen2-1.5B, Qwen2-7B, Qwen2-57B-A14B (Mixture-of-Experts), and Qwen2-72B. All models now feature Group Query Attention (GQA) for faster inference and lower memory usage. Small models (0.5B and 1.5B) use tying embedding to reduce parameter overhead. The base models are pretrained on 32K token contexts with strong extrapolation, while the 7B-Instruct and 72B-Instruct models support up to 128K tokens when augmented with YARN. Training data spans 27 additional languages beyond English and Chinese, including Western and Eastern European, Middle Eastern, Asian, and South Asian languages, with explicit handling of code-switching. Qwen2 achieves state-of-the-art performance in benchmarks, with significant improvements in coding and mathematics over Qwen1.5. The models are released on Hugging Face and ModelScope under an open-source license.

Key Features

Five model sizes: 0.5B, 1.5B, 7B, 57B-A14B (MoE), 72B
Group Query Attention (GQA) on all model sizes for faster inference
Extended context length up to 128K tokens for 7B-Instruct and 72B-Instruct
Multilingual support trained on 27 additional languages beyond English and Chinese
Significantly improved performance in coding and mathematics
Open-source release on Hugging Face and ModelScope
Both base and instruction-tuned models available
Tying embedding applied to small models (0.5B, 1.5B) for parameter efficiency

Pros & Cons

Pros
  • State-of-the-art performance across multiple benchmarks
  • Open-source and free to use under a permissive license
  • Wide range of model sizes for different compute budgets
  • Efficient inference with Group Query Attention
  • Excellent multilingual capabilities covering 27 languages
  • Long context window support (128K tokens) for advanced use cases
  • Active community and ecosystem (Hugging Face, ModelScope, Kaggle)
Cons
  • Context length of 128K only validated for 7B-Instruct and 72B-Instruct models
  • Larger models (57B, 72B) require significant computational resources for inference
  • Multilingual support, while broad, does not cover all languages worldwide
  • Ecosystem and third-party tooling less mature compared to proprietary models like GPT-4

Best For

General language understanding and generationMultilingual applications including translation and code-switchingCoding assistance and mathematical problem solvingLong-document comprehension and summarization (up to 128K tokens)Chatbots and conversational AI with instruction-tuned modelsResearch and experimentation with various model scales

FAQ

What model sizes are available in Qwen2?
Qwen2 offers five model sizes: 0.5B, 1.5B, 7B, 57B-A14B (Mixture-of-Experts), and 72B parameters.
Does Qwen2 support long context?
Yes. The base models are pretrained on 32K tokens and show good extrapolation. The instruction-tuned Qwen2-7B-Instruct and Qwen2-72B-Instruct support up to 128K tokens when augmented with YARN.
What languages does Qwen2 support?
In addition to English and Chinese, Qwen2 is trained on 27 additional languages including German, French, Spanish, Portuguese, Italian, Dutch, Russian, Czech, Polish, Arabic, Persian, Hebrew, Turkish, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Lao, Burmese, Cebuano, Khmer, Tagalog, Hindi, Bengali, and Urdu.
Is Qwen2 open source?
Yes, Qwen2 models are open-sourced and freely available on Hugging Face and ModelScope.
How does Qwen2 compare to Qwen1.5?
Qwen2 brings significant improvements over Qwen1.5, including Group Query Attention on all sizes, extended context length, better multilingual support, and state-of-the-art performance, especially in coding and mathematics.