Ling-mini-2.0

View as Markdown

Ling 2.0 is the Mixture-of-Experts LLM family from inclusionAI (Ant Group), released under the bailing_moe HF architecture (BailingMoeV2ForCausalLM). The line spans a 16 B mini through a 1 T flagship while sharing the same architecture.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Ling-mini-2.0

From the repository root, run LoRA fine-tuning:

automodel examples/llm_finetune/ling/ling_mini_2_0_squad.yaml --nproc-per-node 1

A single 80 GB H100 / A100 fits Ling-mini-2.0 in bf16 with the LoRA defaults in the example. Set distributed.ep_size > 1 for multi-GPU expert parallelism on the larger variants.

Choose a Workflow

GoalStart Here
Low-rank adaptation (LoRA) SFT - Ling-mini-2.0 on SQuADUse ling_mini_2_0_squad.yaml. Minimum hardware: 2x H100 80GB.
Low-rank adaptation (LoRA) SFT - Ling-mini-2.0 on HellaSwagUse ling_mini_2_0_hellaswag.yaml. Minimum hardware: 2x H100 80GB.
Full SFT - Ling-mini-2.0 on HellaSwag, FSDP2 + EP=8Use ling_mini_2_0_sft.yaml. Minimum hardware: 8x H100 80GB.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE)
ArchitectureBailingMoeV2ForCausalLM
Parameters16B total
Hugging Face OrganizationinclusionAI
  • BailingMoeV2ForCausalLM (HF model_type: "bailing_moe")
  • GQA attention; use_qk_norm: true
  • Half RoPE (partial_rotary_factor=0.5)
  • DeepSeek-V3-style routing: sigmoid scoring, per-expert bias, grouped top-k (n_group=8, topk_group=4)
  • 1 shared expert at moe_intermediate_size
  • first_k_dense_replace dense MLP layer(s) at the start of the stack

Available Models

ModelHF ID
Ling-mini-2.0inclusionAI/Ling-mini-2.0