Moonlight-16B-A3B

View as Markdown

Moonlight-16B-A3B is a Mixture-of-Experts (MoE) language model from Moonshot AI with 16B total parameters and 3B active parameters. Moonshot AI reports training it on 5.7T tokens with the Muon optimizer.

The documented workflows cover HellaSwag fine-tuning, packed sequences, pretraining from configuration, and Tulu-3 fine-tuning.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Moonlight

From the repository root, run:

uv run automodel examples/llm_finetune/moonlight/moonlight_16b_te.yaml \
--nproc-per-node 8 \
--checkpoint.enabled true \
--checkpoint.checkpoint_dir checkpoints/moonlight

Choose a Workflow

GoalStart Here
Fine-tune the base checkpointUse moonlight_16b_te.yaml. It fine-tunes on HellaSwag for two epochs with FSDP2 and eight-way expert parallelism.
Fine-tune with packed sequencesUse moonlight_16b_te_packed_sequence.yaml. It packs sequences to 1,024 tokens and uses torch_mm experts.
Pretrain from configurationCustomize the pretraining template. Replace the data paths, configure Slurm, and submit it on one 8-GPU node.
Reproduce the Tulu-3 workflowFollow the Tulu-3 run instructions. Prefilter the data to 2,048 tokens and launch the fine-tuning script with torchrun.

Model Reference

Model Architecture

PropertyValue
Hugging Face ArchitectureDeepseekV3ForCausalLM
Parameters16B total / 3B active
Decoder Layers27
Hidden Size2,048
Attention16 attention heads and 16 key-value heads
Feed-Forward Sizes11,264 dense / 1,408 per expert
Experts64 routed / 2 shared / 6 selected per token
Context Length8,192 tokens
Vocabulary Size163,840
Weight Dtypebfloat16
Upstream Training5.7T tokens with Muon

Available Models

CheckpointType
moonshotai/Moonlight-16B-A3BPretrained base model used by the recipes above
moonshotai/Moonlight-16B-A3B-InstructUpstream instruction-tuned checkpoint