DeepSeek-V3 Technical Report logo

DeepSeek-V3 Technical Report

Free

Strong MoE language model with 671B parameters and 37B activated per token

FreeFree tier
Inputs: textOutputs: text
Type
Open Source
Company
DeepSeek-AI

About DeepSeek-V3 Technical Report

DeepSeek-V3 is a powerful Mixture-of-Experts (MoE) language model developed by DeepSeek-AI, featuring 671 billion total parameters with 37 billion activated per token. It builds upon the architectures validated in DeepSeek-V2, employing Multi-head Latent Attention (MLA) and DeepSeekMoE to achieve efficient inference and cost-effective training. The model introduces an auxiliary-loss-free strategy for load balancing and a multi-token prediction training objective to enhance performance. Pre-trained on 14.8 trillion tokens of diverse, high-quality data, it undergoes supervised fine-tuning and reinforcement learning to fully harness its capabilities. Comprehensive evaluations show that DeepSeek-V3 outperforms other open-source models and rivals leading closed-source models across a wide range of benchmarks.

Key Features

Mixture-of-Experts (MoE) architecture with 671B total parameters, 37B activated per token
Multi-head Latent Attention (MLA) for efficient inference
DeepSeekMoE architecture for cost-effective training
Auxiliary-loss-free strategy for load balancing
Multi-token prediction training objective for stronger performance
Pre-trained on 14.8 trillion diverse and high-quality tokens
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) stages
Outperforms other open-source models and achieves performance comparable to leading closed-source models

Pros & Cons

Pros
  • State-of-the-art performance among open-source models, matching closed-source competitors
  • Efficient inference via MoE with only 37B activated parameters
  • Cost-effective training due to architectural innovations
  • Novel load balancing strategy without auxiliary losses
  • Multi-token prediction objective enhances learning effectiveness

Best For

Large-scale natural language understanding and generationText completion and reasoning tasksQuestion answering and conversational AICode generation and programming assistanceResearch and development in AI language modeling

FAQ

What is the total number of parameters in DeepSeek-V3?
DeepSeek-V3 has 671 billion total parameters, with 37 billion activated per token.
How was DeepSeek-V3 trained?
It was pre-trained on 14.8 trillion diverse and high-quality tokens, followed by supervised fine-tuning and reinforcement learning stages.
What architectures does DeepSeek-V3 use?
It uses Multi-head Latent Attention (MLA) and DeepSeekMoE, which were thoroughly validated in DeepSeek-V2.
How does DeepSeek-V3 compare to other models?
DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models.