DeepSeek-V3

View as Markdown

DeepSeek-V3 is a large-scale Mixture-of-Experts model with 671B total parameters and 37B activated per token. It features Multi-head Latent Attention (MLA), innovative load balancing, and Multi-Token Prediction (MTP). DeepSeek-V3.2 is an updated release with further improvements.

Moonlight by Moonshot AI also uses this architecture with 16B total / 3B activated parameters.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE)
ArchitectureDeepseekV3ForCausalLM / DeepseekV32ForCausalLM
Parameters671B total / 37B active
Hugging Face Organizationdeepseek-ai

Available Models

ModelHF ID
DeepSeek-V3deepseek-ai/DeepSeek-V3
DeepSeek-V3-Basedeepseek-ai/DeepSeek-V3-Base