DeepSeek-V3
DeepSeek-V3
DeepSeek-V3 is a large-scale Mixture-of-Experts model with 671B total parameters and 37B activated per token. It features Multi-head Latent Attention (MLA), innovative load balancing, and Multi-Token Prediction (MTP). DeepSeek-V3.2 is an updated release with further improvements.
Moonlight by Moonshot AI also uses this architecture with 16B total / 3B activated parameters.
Set up NeMo AutoModel with the latest container or follow the installation instructions.