Moonlight-16B-A3B
Moonlight-16B-A3B
Moonlight-16B-A3B is a Mixture-of-Experts (MoE) language model from Moonshot AI with 16B total parameters and 3B active parameters. Moonshot AI reports training it on 5.7T tokens with the Muon optimizer.
The documented workflows cover HellaSwag fine-tuning, packed sequences, pretraining from configuration, and Tulu-3 fine-tuning.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Moonlight
From the repository root, run: